A look back on 2019
The talks that shaped AI Council 2019.
2019 Featured Talks
Highlights from AI Council 2019 — the talks that defined the year.
All 2019 Talks
Every session from 2019 — filter by topic, speaker, or company.
When Testing in Production is a Good Idea
Our business needs us to deliver big improvements to our analytics infrastructure and our client SDKs, on a reliable cadence and with low tolerance for regressions. This session covers some techniques we've developed to use the entropy of production to make this possible. I’ll provide a methodology that developers can use to do the same, especially in contexts with a lot of variability in how their software is used. With these ideas, we’ve been able to make predictable 20% to 40% improvements to the speed of our analysis infrastructure every quarter, with a team of two engineers.

Time Series Prediction with TensorFlow
RNNs and LSTMs have enjoyed great success in text generation algorithms, but their use in other fields has not been as widely studied. We will discuss our experiences and progress using Recurrent Neural Networks to make predictions on arbitrary multivariate time series data. Our first study used weather data from the JFK terminal over several years using the TensorFlow framework. We will discuss the issues related to tuning and validating this model, as well as how we migrated this model into the Model Asset Exchange, which is an IBM hosted API for making predictions on data using pre-trained neural network models. Our insight into tuning this model allowed us to provide another API via Watson Machine Learning, which is a hosted service that allows user defined data and models to be uploaded, trained, and tuned on GPU accelerated on demand hardware using simple remote API calls. We will discuss examples from the financial sector, weather prediction, and other important time series prediction use cases.

Transfer Learning in NLP - How to Help Small Teams Account for Small Datasets
Machine Learning is all about data. Larger companies have been collecting from various sources for years, and are able to build powerful ML models from that data. But what can you do when your own dataset is lacking in size? Transfer Learning has been on the scene for years in Computer Vision, but is just now making a significant impact in NLP. This talk outlines how smaller teams can make efficient use of small, domain specific datasets by utilizing pre-existing models trained on large, public corpuses. Ryan will begin with a summary of Wootric’s journey navigating their problem space, while discussing personal examples of pros and cons to various solutions. He will then describe how transfer learning can be used effectively in NLP, and why having a smaller dataset does not necessarily lead to building an inferior model.

The history and anatomy of Apache Superset
Open source projects can become complex living organisms that grow and bring large communities together. As Apache Superset grows to be 4 year old, we look back on the journey that brought the project to where it is today. More generally, we'll explore what it takes to grow an open source project, a community and a movement. We'll look at a retrospective of the design decisions, technology choices and engineering challenges that have shaped Superset. We'll also take a deep look into the current challenges the community is currently facing, and peak at what is ahead.

Swimming in the Data River, or, when “Streaming Analytics” isn’t
The dirty secret of most “streaming analytics” technologies is that they are just stream processors: they sit on a stream and continuously compute the results of a particular query. They’re good for alerting, keeping a dashboard up-to-date in real time, and streaming ETL, but they’re not good at powering apps that give you true insight into what is happening: for this you need the ability to explore, slice/dice, drill down, and search into the data.This talk will cover the current state of the streaming analytics world, what Druid brings to the table, and some of the technical details behind its design and its integration with Kafka.

Tactical Data Engineering
How do you organize your data so that your users get the right answers at the right time? That question is a pretty good definition of data engineering — but it is also describes the purpose of every DBMS (database management system). And it’s not a coincidence that these are so similar. This talk looks at the patterns that reoccur throughout data management — such as caching, partitioning, sorting, and derived data sets. As the speaker is the author of Apache Calcite, we first look at these patterns through the lens of Relational Algebra and DBMS architecture. But then we apply these patterns to the modern data pipeline, ETL and analytics. As a case study, we look at how Looker’s “derived tables” blur the line between ETL and caching, and leverage the power of cloud databases.

Split Learning: A Resource Efficient Distributed Deep Learning Method without Sensitive Data Sharing
Collaboration in health is heavily impeded by lack of trust, data sharing regulations and limited consent of patients. In settings where different institutions hold different modalities of patient data in the form of electronic health records (EHR), picture archiving and communication systems (PACS) for radiology and other imaging data, pathology test results, or other sensitive data such as genetic markers for disease, collaborative training of distributed machine learning models without any data sharing or leakage of patterns about raw data is desired. In addition the solution needs to be resource efficient in terms of communication bandwidth, computations and memory. This talk is primarily about a recently developed, highly resource efficient method called 'Split Learning' for this very purpose by allowing to perform distributed deep learning under these constraints.

Spatial Data Science Methods for Improving Models
Spatial data science uses many of the same techniques and algorithms as traditional data science, but the spatial component can add a large amount of additional information by combining with other sources at the same location (e.g., census, geolocated tweets), using realtime routing services, or even by using the spatial structure of the distribution of the data. In this talk, I will present lessons learned on extracting more information from spatial data than is typically used in data science projects. I will do this by highlighting two tools we recently used for client projects (spatially-constrained clustering, probabilistic principal component analysis), and present about the structure of spatial data in general that can be readily added to models.

Scaling the best healthcare to everyone, with AI
We are in the brink of a revolution in healthcare. A significant driving force is the recent advancements in Artificial intelligence, especially at the intersection of deep-learning and healthcare. We envision a healthcare system with AI-in-the-loop that is poised to redefine and elevate the role of our doctors - empowering them to deliver the optimal care, when, where and to whom it is most needed. The talk will provide an overview of advancements in applications of deep learning to healthcare. We will also share insights into scaling these academic advancements to a real user-doctor facing system - we will showcase our research at Curai to highlight challenges and practical considerations to scale healthcare access to everyone.

Scaling model training: from flexible training APIs to resource management with Kubernetes
Model training can often be a manual process using notebooks or command line scripts run on a shared server or even a laptop. This is convenient for building intuition, but at some point fails to scale: notebooks and command line scripts generally aren't reproducible, which can lead to confusion about what was running in production when. Similarly, as a machine learning application benefits from an increasing count of models (e.g. a common pattern is developing user-specific models as well as a generic model) or increasingly large datasets, simple tasks like keeping track of training runs and managing computational resources quickly become untenable manually. To help solve these problems, we built an easy-to-use API (that we call Railyard) for training machine learning models, allowing fast, reliable iteration on model training. The Railyard workflow provides an API contract for users. Railyard will fetch your features and labels, split the data into training and test sets, pass along any extra JSON you passed to the API, and handle serialization and evaluation for your fitted estimator, completing your job. Railyard is a Scala service that exposes JSON endpoints for training models and fetching the results of model training runs. The service kicks off the training jobs and performs job-state management to track what is being trained and when the jobs kick off and finish. Railyard uses Kubernetes as an execution engine for all of the model training runs; Kubernetes performs resource allocation and management. This allows us to flexibly support different resource types for training runs requiring, e.g. more memory or GPUs. The combination of flexible API and execution engine facilitates continuous retraining of thousands of models every week, allowing us to quickly evolve machine learning models especially for adversarial machine learning applications like fraud, where models degrade more quickly. As part of continuous retraining, we can not only evaluate individual models, but also more sophisticated compositions of models. In this talk, I'll describe the lessons we learned from building and evolving the Railyard API to support heterogeneous production model training workflows to support production models from logistic regression to deep learning and scaling model training using Kubernetes.

Showing 10 of 41
The voices that shaped 2019
Learn from the engineers at OpenAI, NVIDIA, and Anthropic who are moving the industry forward.

Julian Hyde
Senior Staff Engineer, Google

Abe Gong
Co-Founder & CEO, Great Expectations

Alex Ratner
Author of Snorkel, Stanford University

Ali Hamidi
Lead Data Engineer at Heroku, Salesforce

Amit Ramesh
Software Engineer, Yelp

Andrew Colombi
Co-Founder & CTO, Tonic

Andrew Hoh
Product Manager, Airbnb

Andy Eschbacher
Senior Data Scientist, Carto
Supported by leaders in AI infrastructure





















Voices from 2019


AIC provides an intimate setting for interacting with other folks in the industry, whereas other conferences you may not know anyone you meet in the hallways.
Ryan Boyd, Co-Founder, MotherDuck









