A look back on 2019
The talks that shaped AI Council 2019.
2019 Featured Talks
Highlights from AI Council 2019 — the talks that defined the year.

VEA: Validating, Evolving & Anonymizing Data in Real Time
Albert Franzi Cros · Alpha Health

Why We Defined a Metalanguage for SQL
Lewis Hemens · Dataform

Uncertainty-Aware Food Recognition by Deep Learning
Petia Radeva · University of Barcelona

The Case for Metadata for Machine Learning Platforms
Joerg Schad · ArangoDB
All 2019 Talks
Every session from 2019 — filter by topic, speaker, or company.
VEA: Validating, Evolving & Anonymizing Data in Real Time
At Alpha Health we are developing genuine data products and hence we regard data as one of our most valuable assets. For such products that are in constant development, data also evolves continuously to meet the requirement. Hence, from the product version to version, some data entities are expected to change and even to look completely different. To keep track of the changes over time while improving the understanding of our data, we need to accurately define our data entities using schemas and how these schemas evolve from version to version. This talk aims to cover the integration of JSON-schemas in our data flow in real-time, the benefits of using VEA (Validating, Evolving & Anonymizing) and how this approach can empower the whole company to bring data to the next level. We will focus on how to detect and put invalid data into quarantine, how we evolve data into its latest schema version in a streamlined manner and how we generate de-identified, GDPR-compliant data. We will go over the multiple benefits and challenges we found during the implementation from managing data models, to iterating the infrastructure.

Why We Defined a Metalanguage for SQL
SQL adoption is growing quickly thanks to modern SQL warehouses and SQL technologies, enabling teams to do more and more with their data without having to write and maintain imperative code. Despite this, SQL still has many limitations which make it hard to adopt software engineering best practices such as unit testing, assertions, releases, and writing modular, maintainable code bases. In this talk we take a look at what metaprogramming is and how it can help us build on top of SQL to address these shortcomings, as well as unlock a number of new and interesting use cases. We will walk through the metalanguage and framework we have developed at Dataform, the challenges we faced, the design decisions we made, and many real world examples of things you probably never thought you could (or should) do with SQL.

Uncertainty-Aware Food Recognition by Deep Learning
Recent computer vision approaches assisted by Deep Learning (DL) techniques have shown unexpected advancements in solving problems that did not seem obvious to automatize – such as face recognition, lip reading, plate recognition, or lung cancer diagnosis. However, there has been limited efforts from Computer Vision community to tackle food image recognition. Food image recognition is a big challenge for DL from many points of view: just consider how many dishes exist around the world; or how many names a dish can have! Do we have DL techniques to overcome this? In this talk, we review the field of food image analysis within a new framework: uncertainty-aware multi-task food recognition. After discussing our methodology to advance in this direction, we comment potential research, as well as its social and economic impact. We explain that the DL and Computer Vision community can bring powerful tools for professionals and individuals to get aware not only of what people eat but also of how people eat; which undoubtedly will be a step forward towards better health and well-being. References/Papers: 1. Eduardo Aguilar, Marc Bolaños, Petia Radeva: Regularized uncertainty-based multi-task learning model for food analysis. J. Visual Communication and Image Representation 60: 360-370 (2019) 2. Md. Mostafa Kamal Sarker, Hatem A. Rashwan, Farhan Akram, Estefanía Talavera, Syeda Furruka Banu, Petia Radeva, Domenec Puig: Recognizing Food Places in Egocentric Photo-Streams Using Multi-Scale Atrous Convolutional Networks and Self-Attention Mechanism. IEEE Access 7: 39069-39082(2019) 3. Eduardo Aguilar, Beatriz Remeseiro, Marc Bolaños, Petia Radeva: Grab, Pay, and Eat: Semantic Food Detection for Smart Restaurants. IEEE Trans. Multimedia 20(12): 3266-3275 (2018) 4. Aina Ferrà, Eduardo Aguilar, Petia Radeva: Multiple Wavelet Pooling for CNNs. ECCV Workshops (4) 2018: 671-675 5.Talavera Martinez E, Leyva-Vallina M, Sarker MK, Puig D, Petkov N, Radeva P. "Hierarchical approach to classify food scenes in egocentric photo-streams.", IEEE J Biomed Health Inform. 2019 Jun 12. doi: 10.1109/JBHI.2019.2922390

The Case for Metadata for Machine Learning Platforms
It is well known that data quality and quantity are crucial for building Machine Learning models, especially when dealing with Deep Learning and Neural Networks. But besides the data required to build the model itself, there is another often overlooked type of data required to build a production grade Machine Learning Platform: metadata. Modern Machine Learning platforms contain a number of different components: Distributed Training, Jupyter Notebooks, CI/CD, Hyperparameter Optimization, Feature stores, and many more. Most of these components have associated metadata including versioned datasets, versioned Jupyter Notebooks, training parameters, test/training accuracy of a trained model, versioned features, and statistics from model serving. For the dataops team managing such production platforms, it is critical to have a common view across all this metadata, as we have to ask questions such as: Which Jupyter Notebook has been used to build Model XYZ currently running in production? If there is new data for a given dataset, which models (currently serving in production) have to be updated? In this talk, we look at existing implementations, in particular MLMD as part of the TensorFlow ecosystem. Further, we propose a first draft of a (MLMD compatible) universal Metadata API. We demo the first implementation of this API using ArangoDB.

The Value of Data for Citizens and Health Professionals
Managing and analyzing data is crucial for decision-making, and this also holds true for health policy. This englobes data related to health outcomes, as well as data on patient experience; with the ultimate goal of supporting a health system that is sustainable, focused on people, public, equitable, and adapted to a permanently evolving society.
Talking Bayes to Business: A/B Testing Use Case
Bayesian tools for statistical analysis have made huge leaps forward in the past years, and are on their way to becoming an industry standard. On the other hand, as machine learning tools become more and more prevalent, many issues have come to light: dealing properly with uncertainty, understanding causality, the concept of "statistical significance", etc. In this talk I'll walk you through a basic A/B testing use case and demonstrate the strength challenges of Bayesian tools in dealing with these issues.

Stream Processing Beyond Streaming Data
Stream processing is becoming something like a ""grand unifying paradigm"" for data processing. Outgrowing its original space of real-time data processing, stream processing is becoming a technology that offers new approaches to data processing (including batch processing), real-time applications, and even distributed transactions. We will take a look at these developments from the view of Apache Flink and present some of the major efforts in the Flink community to build a unified stream processor data processing and data-driven applications. Flink already powers many of the world's most demanding stream processing applications. We present the approach of Flink's next generation streaming runtime that also offers a state-of-the-art batch processing experience and performance. A new Machine Learning library, built on top of a unique new API supports many algorithms to train dynamically across static and real-time data. Finally, we look at new building blocks stream processing offers for data-driven applications that open a new direction to solve application consistency. With use cases from different users, we show how companies apply this broader streaming paradigm in practice.

Rethinking Transportation in Cities: Making Traffic Smarter Through Optimization and Location Intelligence
Car traffic is one of the main sources of pollution in cities. In addition, the infrastructure required to absorb all these cars takes up masses of space that could be used for parks, bikes, or simply wider sidewalks. Taxis and rideshare vehicles in particular spend most of their time driving empty. Knowing that this service is essential in cities, what if we could rethink the way this service is provided to make more livable cities? In this talk, I will show how traditional Optimization techniques can be combined with Spatial Data Science to build a powerful algorithm that reduces taxi empty driving hours while maintaining a good level of service in terms of response time. In order to build this algorithm, we start from a basic greedy algorithm with limited use of spatial information. The complexity of the algorithm is gradually increased by introducing linear optimization and the use of third-party footfall data to match supply and demand. Lastly, I will suggest further tips to improve results, as well as client and driver experience. Intermediate and final results are visualized in vector maps, allowing data scientists and users to easily understand supply and demand patterns, and identify improvements on the designed algorithm.

Murron: Reliable Logging Pipeline
Slack is a communication and collaboration platform for teams. Our millions of users spend more than 10 hours connected to the service on a typical working day. Logs & events are the critical components of business applications. The logs help us to understand how a system works, debug, analyze performance, and improve operational efficiency. They are the foundation of building distributed systems. Logging infrastructure is a critical component for Slack; our logging pipeline drives customer billing, usage pattern, and performance analysis of the business-critical systems. This talk walks through the first-generation logging infrastructure and some of the problems we have encountered. We will then cover the second-generation logging infrastructure design, how we added reliability as a built-in feature of the systems, the overall design principles, and some of the best practices for designing scalable logging infrastructure.

Machine Learning for Brain Health and Understanding at Starlab Neuroscience
The convergent development of different technologies is bringing the understanding of the brain, both on healthy and pathological condition, further than ever before. The confluence of Wearables, Neurotechnologies, Augmented and Virtual Reality, Serious Gaming, Data Science, Machine Learning and Artificial Intelligence are a game changer in the way we study brain functionality, make use of it for interacting with the environment, and treat mental and neurological disease. The talk will deal with the combination of Neurotechnologies, Machine Learning and Artificial Intelligence in different Digital Brain Health applications developed at Starlab Neuroscience. Digital markers of brain function will lead in the near future to improved diagnostic, drug discovery, risk analysis, and interactivity. We will show developed methodologies for: stratified performance evaluation of classifiers in operational conditions for Parkinsons’ risk assessment, differential diagnosis in ADHD based on Reservoir Computing, and new treatment outcome prediction in Coma patients. I will go over the technical challenges we faced to develop these applications, but also over some insights that influence the applicability of pure academic data science in the real world.

Showing 10 of 31
The voices that shaped 2019
Learn from the engineers at OpenAI, NVIDIA, and Anthropic who are moving the industry forward.

Agata Lapedriza
Tenured Professor, UOC / Visiting Researcher, MIT Media Lab, UOC / MIT Media Lab

Albert Franzi Cros
Data Engineer, Alpha Health

Alessandro Pregnolato
Head of Data, Marfeel

Alexander Kudryashov
Senior Software Engineer, New Relic

Ananth Packkildurai
Senior Software Engineer, Slack

Andrea Spina
Head of R&D, Radicalbit

Arnau Tibau Puig
Head of Data Science, Letgo

Aureli Soria-Frisch
Director of Neuroscience, Starlab Consulting Division
Supported by leaders in AI infrastructure





















Voices from 2019


AIC provides an intimate setting for interacting with other folks in the industry, whereas other conferences you may not know anyone you meet in the hallways.
Ryan Boyd, Co-Founder, MotherDuck










