A look back on 2019
The talks that shaped AI Council 2019.
2019 Featured Talks
Highlights from AI Council 2019 — the talks that defined the year.
All 2019 Talks
Every session from 2019 — filter by topic, speaker, or company.
Translating Source Code into Natural Language with AI
Software engineering collaboration is hard. Software engineers spend more than 70% of their time learning about their own team’s source code. Enormous size, constant change and intricate dependencies of the source code are the main factors at play. While one software engineer is not responsible for all lines of code for their company, a single line of their code can break the entire company’s app. To help software engineers, we translate source code into natural language to make it easier to search, navigate and understand. At Quod AI, we are building an AI knowledge assistant which generate documentation (in Q&A format) from raw source code. In order to do that, we use neural network models, natural language processing algorithms and statistical models. We retrieve, store and analyze the source code and its history to get insights from the evolution of the code. In this talk we will share some of the insights that we gained from analyzing more than 300 millions lines of code.

Taking Recommendation to the Masses
Recent decades have witnessed a great proliferation of recommendation systems. The technology has brought significant benefits to many business verticals. From earlier algorithms such as similarity-based collaborative filtering to the latest deep neural network-based methods, recommendation technologies have evolved dramatically. To a certain extent, this makes it challenging for practitioners to select and customize the optimal algorithms for a specific business scenario. In addition, operations such as data pre-processing, model evaluation and system operationalization play a significant role in the lifecycle of developing a recommendation system; however, they are often neglected by practitioners. Based on extensive experience in productization of recommendation systems in a variety of real-world application domains, this talk will review and demonstrate the main key tasks in building recommendation systems. It will present best practices and provide examples of democratizing recommendation systems for every organization and the wider community. Open source GitHub repository Microsoft/Recommenders (https://github.com/Microsoft/Recommenders) will be used for the hands-on practice. This repository, where the key topics are shared as Jupyter notebooks and a utility function codebase, is designed to help data scientists quickly grasp basic concepts in a hands-on fashion. It has gained good visibility within the community, with more than 3,600 stars on GitHub. The best practice examples shared in the repository are meant to help developers, scientists or researchers to quickly build production-ready recommendation systems as well as to prototype novel ideas using the provided utility functions. The talk will walk through several recommendation algorithms in order to provide an in-depth understanding of the available techniques.

Sparklens: Understanding the Scalability Limits of Spark Applications
One of the common requests we receive from customers at Qubole is to debug a slow Spark application. Usually this process is done with trial and error, which takes time and requires running clusters beyond normal usage (read wasted resources). Moreover, it doesn’t tell us where to look for further improvements. We at Qubole are looking into making this process more self-serve. Towards this goal we have built Sparklens (https://github.com/qubole/sparklens), an OSS tool based on Spark's event listener framework. From a single run of the application, Sparklens provides insights about scalability limits of a given Spark application. In this talk we will cover what Sparklens does and the theory behind it. We will talk about how the structure of a Spark application puts important constraints on its scalability; how can we find these structural constraints and how to use them as a guide in solving performance and scalability problems of Spark applications. This talk will help the audience with answering the following questions about their Spark applications: 1) Will their application run faster with more executors? 2) How will cluster utilization change as the number of executors changes? 3) What is the absolute minimum time this application will take even if we give it infinite executors? 4) What is the expected wall clock time for the application when we fix the most important structural limits of these applications? Sparklens makes the ROI of additional executors extremely obvious for a given application and needs just a single run of the application to determine how the application will behave with different executor counts. Specifically, it will help managers take the correct side of the tradeoff between spending developer time optimizing applications vs. spending money on compute bills.

Scaling Data Science Teams: Twitter's Perspective
Twitter’s data science organization has grown from 15 people in 2016 to 80 people by the end of this year. In this talk, Miguel Ríos, head of Consumer Data Science at Twitter, will present how they used data and engineering-driven processes to expand their organization globally, the lessons they learned along the way, and the ongoing challenges they face as they prepare for future growth.

Revenue Maximization in the Shared Bike Business Using Network Analysis and Geospatial Mapping
In 2017, Zoomcar launched India's first bike sharing service, PEDL, aimed to make shorter commutes convenient. This business comes with challenges such as maintaining high cycle availability at all times and managing IoT device dysfunctionalities, vandalism, fleet re-balancing and many more.The talk will cover the methodology to overcome these burning challenges, keeping in mind both revenue maximization and customer experience.

Presto: Optimizing Performance of SQL-on-Anything
Presto, an open source distributed SQL engine, is widely recognized for its low-latency queries, high concurrency, and native ability to query multiple data sources. Proven at scale in a variety of use cases at Airbnb, Comcast, GrubHub, Facebook, FINRA, LinkedIn, Lyft, Netflix, Twitter, and Uber, in the last few years Presto experienced an unprecedented growth in popularity in both on-premises and cloud deployments over Object Stores, HDFS, NoSQL and RDBMS data stores. With the ever-growing list of connectors to new data sources such as Azure Blob Storage, Google Cloud Storage, Elasticsearch, Netflix Iceberg, Apache Kudu, and Apache Pulsar, the recently introduced Cost-Based Optimizer in Presto must account for heterogeneous inputs with differing and often incomplete data statistics. This talk will explore this topic in detail as well as discuss best use cases for Presto across several industries. In addition, we will present the recent Presto advancements such as Geospatial analytics at scale and the project roadmap going forward.

Delivering ML Models the Safe and Sane Way
Despite the hype around machine learning and AI, the lifecycle of ML models often end in Kaggle competitions, hackathons and proof of concepts. Very few make it to production because individuals and teams inevitably encounter impediments in deployments, model management, and reproducibility, just to name a few. In this talk, we will share principles and practices on how we can overcome these challenges and enable teams to iteratively deliver ML solutions.

Data Modeling and Processing for a Travel Super App
Traveloka is an app that provides a wide range of travel-related products and services, such as flights, hotels, apartments, theme parks, and even international roaming packages. Having a wide-ranging business makes data modeling particularly challenging: it is like building many data warehouses for different business flows in one place. In order to address that, we developed a modeling method and framework that enables us to model the data across business units, and ensure data is uniform across the board so that data scientists can make sense out of it across all products and services. We developed a data model schema with an inheritance and business glossary concept. The concept enforces uniformity of the data and consistent definitions across all our products. The schema enables data architects to model data schemas, data analysts to describe data definitions, data governance specialists to protect personal data, and data engineers to define cleansing rules, all in one place! The framework is built on top of Python Apache Beam and currently runs on GCP DataFlow. Building on Apache Beam enables us to run the very same framework on our batch and streaming pipeline. The framework is inspired by JSON schema and BigQuery schema. We call it NeoDDL.

Data Architecture 101 for Your Business
Setting up your data architecture can be tricky and confusing without knowing what the future holds for your company’s growth. Some might have attempted to sell you out of the shelf solutions or you could have been overwhelmed by hearing about unlimited different technologies, concepts, big data engines that are scalable without a limit... Right? Or just go with Google Analytics since your marketing team is already keen on that? Do you have a hunch of what you should use? I have worked and built multiple data architectures for companies with different sizes from only few thousands to billions of active users; and used all modern technologies such as Azure SQL Data Warehouse, Redshift, Presto, Hive, Spark, Airflow, Kinesis Data Firehose. From my experiences at Facebook and Microsoft, I know how these tools can be used efficiently and what are the best practices of the industry. In this talk I will guide you what solutions are available for all company sizes, when is the right time to add or replace architecture elements for better scaling and/or better engineering. What are the caveats and deep technical tricks to get the most out of these tools. Moreover, I will answer how to avoid building or setting up overcomplicated systems, and when should you hire data scientists or data engineers.

Causal Inference: Making the Right Intervention
Consider an organization seeking to improve their operations, using their historical data. During this type of analysis the commonly known fact that “correlation does not imply causation” comes to life. It is crucial to distinguish between events that *cause* existing inefficiencies and those that merely correlate. Spending money to fix something that is not the root cause of the problem could be an expensive folly. Causal inference aims to determine which available controls drive specific outcomes. This is a distinctly more demanding condition than learning the correlation. Many machine learning approaches disregard causal inference, despite a wide range of approaches to causal inference having been proposed in the literature. This talk will discuss the importance of causal models, as well as some of the most state-of-the-art methods for reasoning.

Showing 10 of 16
The voices that shaped 2019
Learn from the engineers at OpenAI, NVIDIA, and Anthropic who are moving the industry forward.

Arpit Agarwal
Director of Data Science, Zoomcar

Ashish Dubey
VP of International Solution Architects, Qubole

Bence Faludi
Independent Consultant

Bin Fan
Founding Member of Engineering, Alluxio Inc

Chris Hausler
Senior Manager, Data Science, Zendesk

David Tan
Software Engineer, Thoughtworks

Greg Roodt
Data Engineering Lead, Canva

Kamil Bajda-Pawlikowski
CTO, Starburst
Supported by leaders in AI infrastructure





















Voices from 2019


AIC provides an intimate setting for interacting with other folks in the industry, whereas other conferences you may not know anyone you meet in the hallways.
Ryan Boyd, Co-Founder, MotherDuck









