A look back on 2017
The talks that shaped AI Council 2017.
2017 Featured Talks
Highlights from AI Council 2017 — the talks that defined the year.
All 2017 Talks
Every session from 2017 — filter by topic, speaker, or company.
Zip codes and other lies your map told you
Your spatial data might be lying to you. Zip code is the most common piece of geo-data analysts and data scientists see, but it has many quirks that can derail your analysis and lead to false conclusions. We'll look at the zip code and learn exactly what it is - and what it isn't. To do this we'll take a look at how a piece of mail get from point A to point B and take very quick trip though the history of the U.S. postal system before looking at other data you can collect that may be more appropriate for spatial purposes than zip code. Then we'll turn our attention to other ways seemingly good spatial data can lie to you: trap streets and paper towns.

You Won't Believe How We Optimize our Headlines
At BuzzFeed we generate hundreds of articles a day, so choosing better headlines can save us from substantial losses in our audience engagement. Our solution is a tool that takes in multiple headline and thumbnail options for an article and decides which combination is most effective. In this talk, I discuss the models that perform best for this tool under different product scenarios. I also discuss causal analysis of the effectiveness of this tool when A/B testing is infeasible.

Worse Case Scenario in the Database
Technical debt in the code base is one thing, but what to do about technical debt in the database? When a production system hasn't been touched in years, the data models can get nasty and restoring order can seem impossible. How do you untangle the mess and restore the database to efficient service?

When Production Machine Learning Fails
Recently, there has been substantial media attention placed on failures in machine learning systems. Here I will present some of the challenges that Predata has faced in building predictive products, as well as giving brief overviews of a few techniques for combatting these. While by no means an exhaustive catalogue of ML failure modes, the nature of our prediction problem and our data has led us to face challenges including but not limited to class imbalance, non-stationarity, seasonality, concept drift, and difficulty establishing good metrics and loss functions.
Video Understanding at Scale: Deep Learning in a Serverless Infrastructure
Video is a complex data structure: it's composed of large amounts of images, sound and text. Not only it's complex but, to understand that data, you need complex Machine Learning models. In this talk, we are going to talk a bit about the work we have done at Uru to understand videos in the media and advertising vertical, talking specifically about how we managed to leverage Deep Learning models to extract meaningful data in those videos and how we leveraged a serverless architecture to achieve real time processing speeds
Using Causal Forests for Subpopulation Identification in Randomized Clinical Trials
Significant scientific problems require a combination of prediction and inference, however, the majority of machine learning techniques are not well-suited for estimating causal effects. Recently, several novel approaches have attempted to combine predictive modeling with causal inference to identify heterogenous treatment effects in observational data — subgroups that may have a significantly different outcome than the population average. This talk will review two recent papers that employed the causal forest approach to estimate subgroup treatment effects in randomized clinical trials. The Systolic Blood Pressure Intervention Trial (SPRINT), compared standard versus intensive systolic blood pressure targets, while the Look AHEAD examined the effects of an intensive diabetes lifestyle intervention on cardiovascular mortality. Reanalyzing these trials using causal forests, we found that in both cases, average treatment effects may have masked important sources of heterogeneity in trial outcomes. These findings bring to questions for future work: First, are there data-driven methods that can objectively identify subgroups for better precision? Second, can we use these large RCTs to build better predictive models on observational data?

Using Apache Spark for processing trillions of records each day at Datadog
Massively scaling Apache Spark can be challenging, but it’s not impossible. In this session we’ll share Datadog’s path to successfully scaling Spark and the pitfalls we encountered along the way. We’ll discuss some low-level features of Spark, Scala, JVM, and the optimizations we had to make in order to scale our pipeline to handle trillions of records every day. We’ll also talk about some of the unexpected behaviors of Spark regarding fault-tolerance and recovery—including the ExternalShuffleService, recomputing partitions, and Shuffle Fetch failures—which can complicate your scaling efforts.

Using Apache Arrow, Calcite and Parquet to build a Relational Cache
Everybody wants to get to data faster. As we move from more general solution to specific optimization techniques, the level of performance impact grows. This talk will discuss how layering in-memory caching, columnar storage and relational caching can combine to provide a substantial improvement in overall data science and analytical workloads. It will include a detailed overview of how you can use Apache Arrow, Calcite and Parquet to achieve multiple magnitudes improvement in performance over what is currently possible. We'll start by talking about in-memory caches and the difference between block-based and data-aware caching strategies. We'll discuss the deployment design of this type of solution as well as cover the strengths of each. There will also be a discussion of the relationship of security and predicate application in these scenarios. Then we'll go into detail about how columnar storage formats can further enhance performance by minimizing read time, optimizing for vectorized in-memory processing and powerful compression techniques. Lastly, we'll introduce a much more advanced way to speed access to data called relational caching. Relational caching builds a cache on columnar in-memory caching techniques but also includes a full comprehension of how data is being used and how different forms of data relate to each other. This will include leveraging multiple sorting and partitioning strategies as well as maintaining multiple related derivations of data for different types of access patterns. As part of this and we also cover approaches to data ttl, relational cache consistency and several different approaches to data mutation and real-time updates.

TimescaleDB: Rearchitecting a SQL database for time-series data
Today everything is instrumented, generating more and more time-series data streams that need to be monitored and analyzed. When it comes to storing this data, many developers start with some well-trusted system like PostgreSQL. But when their data hits a certain scale, they often give up its query power and ecosystem by migrating to some NoSQL or other "modern" time-series architecture. In this talk, I describe why this perceived trade-off isn't necessary, and how we've built an efficient, scalable time-series database engineered up from PostgreSQL. In particular, the nature of time-series workloads one finds in devops, monitoring, IoT, finance, and elsewhere -- inserting new data about recent events -- presents very different demands than general transactional (OLTP) workloads. We've architected our time-series database to take advantage of and embrace these differences. The system architecture automatically partitions data across both time and space, even though it exposes the illusion of a single continuous table -- a hypertable -- across all of your data spread across one or many servers. Its distributed query optimizations both hide the fact that users are interacting with many "chunks" of data, which are right-sized by volume and time constraints, and minimize which and how chunks are accessed to answer queries. In fact, the database supports "full SQL" against this hypertable (e.g., secondary indexes, rich query predicates and group bys, aggregations, windowing functions, upserts, CTEs, JOINs). Through performance benchmarks, I show how the database scales much better than PostgreSQL, even on a single node. In particular, it avoids the "performance cliff" that vanilla PostgreSQL experiences at 10s of millions of rows, while maintaining robust performance past 100B rows. The database is implemented as a PostgreSQL extension, released under the Apache 2 license.

The Trade-off between Strict Validation and Accepting Anything
When building a data pipeline, we need to decide if we should strictly validate incoming data, and discard anything that we don't support, or if we should be flexible, and accept anything so we can analyze it later. In this talk, I'll discuss how the compromise we've reached at Bluecore, where we both record the "raw" data to recover from bugs or mistakes, as well as strictly validated data. I'll talk about why we think that validating up front is the better choice when building data intensive applications.
Showing 10 of 44
The voices that shaped 2017
Learn from the engineers at OpenAI, NVIDIA, and Anthropic who are moving the industry forward.

Julian Hyde
Senior Staff Engineer, Google

Andreas Mueller
Associate Research Scientist, Data Science Institute, Columbia University

Andy Turley
Lead Software Engineer, Wallaroo Labs

Anne Bauer
Senior Data Scientist, The New York Times

Ben Davis
Co-Founder & CTO, Gather

Brunno Attorre
Co-Founder & CTO, URU

Burak Yavuz
Software Engineer, Databricks

Christian Romming
Founder & CEO, ETLeap
Supported by leaders in AI infrastructure





















Voices from 2017


AIC provides an intimate setting for interacting with other folks in the industry, whereas other conferences you may not know anyone you meet in the hallways.
Ryan Boyd, Co-Founder, MotherDuck









