AI Council 2020

A look back on 2020

The talks that shaped AI Council 2020.

2020 San Francisco — Talks

All 2020 Talks

Every session from 2020 — filter by topic, speaker, or company.

· Talk

ML Feature Engineering is a two-part problem: 1. Generating Realtime Features for Serving 2. Generating Historically Realtime Features for training We will briefly introduce how we address the former problem. The primary focus of this talk is about generating Historically Realtime Features for training. One way to think of generating Historically-Real-Time Features is to: 1.) Travel back in time to a particular state of the world, as represented by data in production systems 2.) Snapshot it, and 3.) Compute aggregations over the snapshot. This is a useful visualization to understand the problem, but it is an intractable approach - especially in terms of compute and storage. We will introduce the algorithm in Zipline that makes backfilling features, with Historically Realtime values, feasible. We will borrow a few concepts from Abstract Algebra / Category Theory, but everything will be introduced from first principles.

Zipline - Airbnb's Declarative Feature Engineering Framework
· Talk

ML developers and data scientists are increasingly tasked with extracting value from torrents of visual data (images, videos, etc). However, handling big-visual-data requires expertise in large scale infrastructure and data management solutions were not developed with ML workflows in mind. ApertureData Platform recognizes the unique characteristics of visual data as well as the importance of its associated metadata, streamlining access and extraction operations across large visual data for scalable ML deployments. Our simple ML-aware interface cuts platform engineering time by months. What current makeshift solutions fail to address is that as ML gets commercialized, managing the onslaught of real visual data is going to be a killer for real deployments. Our talk will explain why status quo needs to be challenged, how ApertureData Platform achieves the performance and functionality important for a wide range of visual ML driven application domains, and demonstrate some real world use cases.

Time to Rethink Visual Data Management for Machine Learning
· Talk

Enabling responsible development of artificial intelligent technologies is one of the major challenges we face as the field moves from research to practice. Researchers and practitioners from different disciplines have highlighted the ethical and legal challenges posed by the use of machine learning in many current and future real-world applications.  Now there are calls from across the industry (academia, government, and industry leaders) for technology creators to ensure that AI is used only in ways that benefit people and “to engineer responsibility into the very fabric of the technology.”   Overcoming these challenges and enabling responsible development is essential to ensure a future where AI and machine learning can be widely used. In this talk we will cover two of our open source toolkits InterpretML and Fairlearn. We will discuss their integration with the Azure Machine Learning platform, share lessons learned and future steps.

Responsible AI – Model Interpretability and Fairness
· Talk

Statistics, machine learning and programming are commonly referenced tools of the data science trade. A less often talked about, but entirely indispensable skill for building great data products is knowing what to build and why. In this talk, I will touch upon the increasing convergence between the skills of an effective data scientist and skills that have traditionally fallen in the realm of product management.

The Unreasonable Effectiveness of Product Sense
Keynote
· Keynote

Deep Learning is disrupting decades' worth of data infrastructure building. Specifically, we will look at retrieval problems. These are core to similarity search, deduplication, recommender systems, matching, feed ranking, personalization, and many more. While deep learning can unlock greater relevance and accuracy, it also presents significant operational and scientific challenges. Collecting data (correctly) is far from being trivial because of complex data interactions, training new models requires scale and distribution that most companies are not comfortable using, and real-time serving (efficiently) is almost impossible with tools that exist today. We will deep-dive into how HyperCube is thinking about unifying these applications under a single abstraction and a unified system.

Real-time Retrieval with Deep Learning: Benefits and Challenges
· Talk

At Uber, we have a variety of machine learning applications including matching, pricing, recommendation, and personalization, resulting in a large number of machine learning models to manage in production. Building machine learning models is an iterative process that spans across a set of stages of a lifecycle. Michelangelo Gallery is Uber's scalable model management system to save, serve, observe, and orchestrate the flow of models across different stages of the ML lifecycle at scale. In this session, we will review the machine learning model lifecycle, present the challenges of managing models at scale in production systems, and give lessons learned/design considerations for building scalable production solutions. We will highlight case studies from Uber describing how Gallery has helped automate workflows, improve iteration velocity, and promote best practices across users.

Keynote
· Keynote

As a field, we often hear about success stories. This is true in research, where a publishing incentive can pressure authors to focus on consistently exceeding state of the art results. It is also true in industry, where companies attempt to attract engineering talent by describing how impressive their production ML systems are. However, every practitioner here knows that in engineering and in ML, the road to success is paved with failures. The field of ML in production is new, and so has a lack of cautionary tales of things that can go wrong with models. This talk will try to help correct that. We will discuss challenges such as performance mismatch between offline training and online inference, feature generation and data leakage, and adequate roadmap planning for ML.

Pitfalls and Challenges of ML-Powered Applications
Workshops
· Talk

dbt is an open source data transformation tool maintained by Fishtown Analytics. In this session, Fishtown Analytics product manager, Jeremy Cohen, will introduce attendees to dbt and why it is used (and loved) by thousands of companies. The workshop will include a presentation on the dbt viewpoint and key functionality around data modeling, testing, and documentation. The presentation will be followed by a Q&A and then a working session where attendees can set up their dbt project or workshop specific challenges with Jeremy.

Meet dbt: The Data Transformation Tool Used by JetBlue, GitLab, Wistia, and Away
· Talk

At Netflix, we reach out to our members about new recommendations in the most effective way, through emails and notifications. We deliver billions of messages per year, and we work on the algorithms that decide what to send, when and to whom. The personalized system employs algorithms at the intersection of Personalization, Reinforcement Learning and Causal Inference. In this talk, we will take the audience through a tour of the continuous explore and exploit system we have iterated on, and showcase a few particularly complex and fascinating challenges with offline evaluation and experimentation.

MESA: Building a Personalized Messaging System at Netflix
· Talk

This talk will cover building conversational AI using deep learning technologies and lessons learnt in developing conversational interfaces. The first part of the talk will describe recent advances in deep learning that has led to tremendous progress in natural language processing and is making conversational AI a reality. Conversational AI includes intent classification, sequence labeling, understanding dialogs and context, and coming up with responses to users messages. The second part of the talk will address lessons learned developing conversational virtual agents. A conversational virtual needs to be personable in addressing, adaptive in understanding, and available to automate different supported tasks. Overall, key take aways from the talk will be better understanding of (1) deep learning techniques for natural language processing and (2) interaction patterns for automation using a conversational virtual agents.

Lessons Developing Conversational AI Virtual Agents

Showing 10 of 17

2020 San Francisco — Speakers

The voices that shaped 2020

Learn from the engineers at OpenAI, NVIDIA, and Anthropic who are moving the industry forward.

Mitul Tiwari, Co-founder & CTO, Stealth

Co-founder & CTO, Stealth

Alyssa Ransbury, Security Engineer, Square

Security Engineer, Square

Edo Liberty, Founder, HyperCube

Founder, HyperCube

Emmanuel Ameisen, ML Engineer, Stripe

ML Engineer, Stripe

Grace Huang, Machine Learning Engineering Manager, Netflix

Machine Learning Engineering Manager, Netflix

Haytham Abuelfutuh, Software Engineering Manager, Lyft

Software Engineering Manager, Lyft

Jeremy Cohen, Associate Product Manager, Fishtown Analytics

Associate Product Manager, Fishtown Analytics

Luis Remis, CTO, ApertureData

CTO, ApertureData