AI Council 2017

A look back on 2017

The talks that shaped AI Council 2017.

2017 San Francisco — Talks

All 2017 Talks

Every session from 2017 — filter by topic, speaker, or company.

· Talk

With recent advances in hardware, frameworks, and research, Deep Learning has emerged as an indispensable technique for solving many data science and AI problems over the last few years. Like any tool, however, it is important to understand when and how to apply it, how to frame your problem in a manner that allows you to apply the tool effectively, as well as what decisions and compromises the machine learning practitioner must make to apply the model on production data and in production systems. In this talk, we will present the lessons we’ve learned developing a deep learning model to handle the distinctive problem eBay faces in recommender systems. We will specifically address the following topics: When to use deep learning rather than other kinds of machine learning algorithms How to frame your problem as one that can be optimized for a deep learning model How to select your training data How to design the right evaluation measures for your model Design considerations for taking your deep learning model into production

Why, When, How: Lessons Learned in Applying Deep Learning to Real-World Problems
· Talk

Twitter is all about real-time at scale. To achieve real-time performance, Twitter has developed, deployed and open-sourced Heron, the next-generation cloud streaming engine. The amount of data that need to be processed in Twitter’s data centers changes significantly due to expected and unexpected global events. For example, during the Super Bowl, there are spikes of tweets that all need to be processed in real-time. Similarly, unexpected events such as natural disasters can generate very large volumes of data. In this talk we will describe how Twitter and Microsoft have been collaborating to transform Heron into a truly elastic system that can support dynamic load changes. We'll present how we adapted several components of the system, such as the scheduler and resource manager, to make them able to seamlesly support elasticity. We will also present our future plans for scaling that current and future Heron contributors could tackle.

Twitter Heron: The Path Towards Elastic Streaming
· Talk

With more than a decade of Big Data experience now behind us, we’ll talk with a few veterans from the front lines about what skills mattered — and what didn't — on the first data engineering teams at Silicon Valley’s data-intensive start-ups. This will be a must-watch panel for those looking to build out a data-engineering function or get into the field themselves. We’ll explore the build vs. buy debate, looking at which classes of software teams hand-rolled, borrowed (and extended) from open-source projects, or bought - and discuss how that mix is changing. We’ll learn about the hardest parts of data pipelines and the data stack to build. And we’ll highlight the necessary skills that make data engineers different than data scientists.

The Right Stuff: Lessons Learned from a Decade of Data Engineering
· Talk

The total amount of data available to human beings currently doubles every 18-24 months, giving data scientists an unprecedented opportunity to push further than ever the boundaries of human knowledge. This is an exciting time for data professionals. Many are hopeful that these huge loads of data will enable data-greedy algorithms like deep neural networks to unlock a myriad of new possibilities for humankind. But can big data really answer all our questions? No matter how useful and powerful, in the wrong hands, data can also easily lead to ill-informed decisions and wrong assumptions. In her talk, Jennifer will cover the reasons why better algorithms matter just as much as the amount of data available, and will describe the dangers and perils that the data scientist of the future will need to thwart using increasingly advanced mathematical knowledge, refined strategies and human rationality.

The Limitations of Big Data in Predictive Analytics
· Talk

A/B testing is a well-understood tool for causal inference in web companies. However, it is not a panacea and often fails when sample sizes are small, measurement lags long, and the treatment space that you want to explore large. At Opendoor, we face all of these problems. American homes represent a $25 trillion asset class, with very little liquidity. Selling a home on the market takes months of hassle and uncertainty. Opendoor offers to buy houses from sellers, charging a fee for this service. Opendoor bears the risk in reselling the house and needs to understand the effectiveness of different liquidity models. Key metrics and resale outcomes can take many months to measure, suggesting that A/B testing may not be the best tool. In this talk we'll cover the ingredients of a simulation-based inference -- from how to define a good data-generating process to user models -- and will walk through a case study in residential real estate. We'll discuss how it obviated the need for certain A/B tests and allowed us to become more efficient in designing the necessary ones.

Simulation-based Inference: Advantages Over A/B Testing in Real Estate
· Talk

This talk deep dives into how Facebook managed to convert a gigantic Hive batch processing job that uses 6000 CPU days to run on Spark with 1/4 CPU at 1/4 latency. To accomplish this, we made numerous stability and performance improvements to Apache Spark, tuned configurations and optimized our business logic. Nominated among the Top 10 blog posts of 2016 from Apache Spark, this talk describes the experiences and lessons learned while scaling Spark to replace one of Facebook's Hive workloads. Examples include taking one of the existing pipelines and migrating it to spark to enable fresher feature data, and improve manageability. This led to major realiability improvements including making the PipedRDD more robust to fetch failure gracefully, as well as a less disruptive cluster restart. In addition, performance optimizations were also made as part of the migration to spark such as reducing shuffle write latency which led to a CPU improve of up to 50% for jobs writing a high number of shuffle partitions.

Scaling Up Spark at Facebook – a 60TB Production Use Case
· Talk

Many emerging Big Data problems are in fact "Fast Big Data problems" where data has to be accessed with very low response time. The rise of in-memory platforms like Apache Spark is an indication of this trend. However most existing distributed in memory platforms such as Apache Spark rely on the system layer that isn't built for high performance in-memory processing. In addition, the existing system software doesn’t optimally use the parallelism inherent in the new architectures where we have CPUs with many cores and SSDs with many flash channels. Other issues such as persistence and uniform access to a large memory space in the cluster also need a fundamental rethinking.

Real-time System Computing Engines
· Talk

Solving problems in the real world with machine learning can be challenging: data from real processes doesn't usually behave like the data from example use cases of ML methods! This talk, designed for applied machine learning practitioners (in anything from business to science to social research), will inspire hope: many challenges of real and messy data can be overcome using straightforward analysis of that data! Alyssa will cover a robust way to add error bars to any number of complex metrics, a strategy for monitoring models in production when you can't always observe an outcome, and a way to plainly explain the decisions made by black-box models.

Practical Solutions for Annoying Machine Learning Problems
· Talk

In building data products at scale there exists a spectrum of endeavors, at one end of which is data analysis and model prototyping and at the other end are data engineering pipelines. Tools such as Scikit-learn and Tensorflow have made former accessible while Spark and other big data stacks have addressed needs on the latter end of the spectrum. Somewhere in the middle of this spectrum is the challenge of operationalizing machine learning models. In this talk, we will share practical lessons and patterns for building machine learning (ML) models in production, based on our experience with search ranking and recommendation systems at Instacart. As part of this I will include a detailed discussion on the technical challenges in building a ML features pipeline, one of which is now shared across multiple data products at Instacart.

Practical Lessons for Building Machine Learning Models in Production
· Talk

Coinbase is the one of the largest digital currency exchanges in the world. We store about $1B of digital currency (bitcoin, litecoin, ether) on behalf of our users. Given the instant nature of digital currency and that it can't be revoked, we have one of the hardest payment fraud and security problems in the world. We are hit by the most sophisticated scammers constantly and consequently we are at the forefront of the fight against fraud. We've witnessed and solved loopholes exploited by fraudsters years ahead of the broader industry (e.g., vulnerabilities in second-factor tokens delivered by SMS, phone porting attacks, loopholes in online identity verification, etc.). In this talk, I'll present examples of scammer trends and techniques we've seen through the past years. I'll also talk about our risk program that relies on rules-based systems, supervised and unsupervised machine learning as well as highly-skilled human fraud fighters.

Payment Fraud in Digital Currency

Showing 10 of 26

2017 San Francisco — Speakers

The voices that shaped 2017

Learn from the engineers at OpenAI, NVIDIA, and Anthropic who are moving the industry forward.

Benn Stancil, Founder, Mode

Founder, Mode

Mike Driscoll, Co-Founder & CTO, Rill Data

Co-Founder & CTO, Rill Data

Paul Dix, Founder & CTO, InfluxData

Founder & CTO, InfluxData

Alyssa Frazee, Machine Learning Engineer, Stripe

Machine Learning Engineer, Stripe

Ashvin Agrawal, Senior Research Engineer, Microsoft

Senior Research Engineer, Microsoft

Ben Hamner, Co-founder & CTO, Kaggle

Co-founder & CTO, Kaggle

Chris Hartfield, Sr. Data Platform Engineer, Clover Health

Sr. Data Platform Engineer, Clover Health

Daniel Galron, Research Scientist & Engineer, eBay

Research Scientist & Engineer, eBay

2017 San Francisco — Sponsors

Supported by leaders in AI infrastructure

Snowflake
TextQL
HEX
Databricks
Braintrust
ClickHouse
Snorkel
Datalinks
Airbyte
Render
Turbopuffer
DigitalOcean
CockroachDB
bem
Preset
LanceDB
Chalk
Unstructured
MotherDuck
Crux
TOPK
2017 San Francisco — Testimonials

Voices from 2017

AIC provides an intimate setting for interacting with other folks in the industry, whereas other conferences you may not know anyone you meet in the hallways.
Ryan Boyd, Co-Founder, MotherDuck
Priya Nair at the panel discussion