AI Council 2017

A look back on 2017

The talks that shaped AI Council 2017.

2017 New York — Talks

All 2017 Talks

Every session from 2017 — filter by topic, speaker, or company.

· Talk

Your spatial data might be lying to you. Zip code is the most common piece of geo-data analysts and data scientists see, but it has many quirks that can derail your analysis and lead to false conclusions. We'll look at the zip code and learn exactly what it is - and what it isn't. To do this we'll take a look at how a piece of mail get from point A to point B and take very quick trip though the history of the U.S. postal system before looking at other data you can collect that may be more appropriate for spatial purposes than zip code. Then we'll turn our attention to other ways seemingly good spatial data can lie to you: trap streets and paper towns.

Zip codes and other lies your map told you
· Talk

At BuzzFeed we generate hundreds of articles a day, so choosing better headlines can save us from substantial losses in our audience engagement. Our solution is a tool that takes in multiple headline and thumbnail options for an article and decides which combination is most effective. In this talk, I discuss the models that perform best for this tool under different product scenarios. I also discuss causal analysis of the effectiveness of this tool when A/B testing is infeasible.

You Won't Believe How We Optimize our Headlines
· Talk

Technical debt in the code base is one thing, but what to do about technical debt in the database? When a production system hasn't been touched in years, the data models can get nasty and restoring order can seem impossible. How do you untangle the mess and restore the database to efficient service?

Worse Case Scenario in the Database
· Talk

Recently, there has been substantial media attention placed on failures in machine learning systems. Here I will present some of the challenges that Predata has faced in building predictive products, as well as giving brief overviews of a few techniques for combatting these. While by no means an exhaustive catalogue of ML failure modes, the nature of our prediction problem and our data has led us to face challenges including but not limited to class imbalance, non-stationarity, seasonality, concept drift, and difficulty establishing good metrics and loss functions.

· Talk

Video is a complex data structure: it's composed of large amounts of images, sound and text. Not only it's complex but, to understand that data, you need complex Machine Learning models. In this talk, we are going to talk a bit about the work we have done at Uru to understand videos in the media and advertising vertical, talking specifically about how we managed to leverage Deep Learning models to extract meaningful data in those videos and how we leveraged a serverless architecture to achieve real time processing speeds

· Talk

Significant scientific problems require a combination of prediction and inference, however, the majority of machine learning techniques are not well-suited for estimating causal effects. Recently, several novel approaches have attempted to combine predictive modeling with causal inference to identify heterogenous treatment effects in observational data — subgroups that may have a significantly different outcome than the population average.  This talk will review two recent papers that employed the causal forest approach to estimate subgroup treatment effects in randomized clinical trials. The Systolic Blood Pressure Intervention Trial (SPRINT), compared standard versus intensive systolic blood pressure targets, while the Look AHEAD examined the effects of an intensive diabetes lifestyle intervention on cardiovascular mortality. Reanalyzing these trials using causal forests, we found that in both cases, average treatment effects may have masked important sources of heterogeneity in trial outcomes. These findings bring to questions for future work: First, are there data-driven methods that can objectively identify subgroups for better precision? Second, can we use these large RCTs to build better predictive models on observational data?

Using Causal Forests for Subpopulation Identification in Randomized Clinical Trials
· Talk

Massively scaling Apache Spark can be challenging, but it’s not impossible. In this session we’ll share Datadog’s path to successfully scaling Spark and the pitfalls we encountered along the way. We’ll discuss some low-level features of Spark, Scala, JVM, and the optimizations we had to make in order to scale our pipeline to handle trillions of records every day. We’ll also talk about some of the unexpected behaviors of Spark regarding fault-tolerance and recovery—including the ExternalShuffleService, recomputing partitions, and Shuffle Fetch failures—which can complicate your scaling efforts.

Using Apache Spark for processing trillions of records each day at Datadog
· Talk

Everybody wants to get to data faster. As we move from more general solution to specific optimization techniques, the level of performance impact grows. This talk will discuss how layering in-memory caching, columnar storage and relational caching can combine to provide a substantial improvement in overall data science and analytical workloads. It will include a detailed overview of how you can use Apache Arrow, Calcite and Parquet to achieve multiple magnitudes improvement in performance over what is currently possible. We'll start by talking about in-memory caches and the difference between block-based and data-aware caching strategies. We'll discuss the deployment design of this type of solution as well as cover the strengths of each. There will also be a discussion of the relationship of security and predicate application in these scenarios. Then we'll go into detail about how columnar storage formats can further enhance performance by minimizing read time, optimizing for vectorized in-memory processing and powerful compression techniques. Lastly, we'll introduce a much more advanced way to speed access to data called relational caching. Relational caching builds a cache on columnar in-memory caching techniques but also includes a full comprehension of how data is being used and how different forms of data relate to each other. This will include leveraging multiple sorting and partitioning strategies as well as maintaining multiple related derivations of data for different types of access patterns. As part of this and we also cover approaches to data ttl, relational cache consistency and several different approaches to data mutation and real-time updates.

Using Apache Arrow, Calcite and Parquet to build a Relational Cache
· Talk

Today everything is instrumented, generating more and more time-series data streams that need to be monitored and analyzed. When it comes to storing this data, many developers start with some well-trusted system like PostgreSQL. But when their data hits a certain scale, they often give up its query power and ecosystem by migrating to some NoSQL or other "modern" time-series architecture. In this talk, I describe why this perceived trade-off isn't necessary, and how we've built an efficient, scalable time-series database engineered up from PostgreSQL. In particular, the nature of time-series workloads one finds in devops, monitoring, IoT, finance, and elsewhere -- inserting new data about recent events -- presents very different demands than general transactional (OLTP) workloads. We've architected our time-series database to take advantage of and embrace these differences. The system architecture automatically partitions data across both time and space, even though it exposes the illusion of a single continuous table -- a hypertable -- across all of your data spread across one or many servers. Its distributed query optimizations both hide the fact that users are interacting with many "chunks" of data, which are right-sized by volume and time constraints, and minimize which and how chunks are accessed to answer queries. In fact, the database supports "full SQL" against this hypertable (e.g., secondary indexes, rich query predicates and group bys, aggregations, windowing functions, upserts, CTEs, JOINs). Through performance benchmarks, I show how the database scales much better than PostgreSQL, even on a single node. In particular, it avoids the "performance cliff" that vanilla PostgreSQL experiences at 10s of millions of rows, while maintaining robust performance past 100B rows. The database is implemented as a PostgreSQL extension, released under the Apache 2 license.

TimescaleDB: Rearchitecting a SQL database for time-series data
· Talk

When building a data pipeline, we need to decide if we should strictly validate incoming data, and discard anything that we don't support, or if we should be flexible, and accept anything so we can analyze it later. In this talk, I'll discuss how the compromise we've reached at Bluecore, where we both record the "raw" data to recover from bugs or mistakes, as well as strictly validated data. I'll talk about why we think that validating up front is the better choice when building data intensive applications.

Showing 10 of 44

2017 New York — Speakers

The voices that shaped 2017

Learn from the engineers at OpenAI, NVIDIA, and Anthropic who are moving the industry forward.

Julian Hyde, Senior Staff Engineer, Google

Senior Staff Engineer, Google

Andreas Mueller, Associate Research Scientist, Data Science Institute, Columbia University

Associate Research Scientist, Data Science Institute, Columbia University

Andy Turley, Lead Software Engineer, Wallaroo Labs

Lead Software Engineer, Wallaroo Labs

Anne Bauer, Senior Data Scientist, The New York Times

Senior Data Scientist, The New York Times

Ben Davis, Co-Founder & CTO, Gather

Co-Founder & CTO, Gather

Brunno Attorre, Co-Founder & CTO, URU

Co-Founder & CTO, URU

Burak Yavuz, Software Engineer, Databricks

Software Engineer, Databricks

Christian Romming, Founder & CEO, ETLeap

Founder & CEO, ETLeap

2017 New York — Sponsors

Supported by leaders in AI infrastructure

Snowflake
TextQL
HEX
Databricks
Braintrust
ClickHouse
Snorkel
Datalinks
Airbyte
Render
Turbopuffer
DigitalOcean
CockroachDB
bem
Preset
LanceDB
Chalk
Unstructured
MotherDuck
Crux
TOPK
2017 New York — Testimonials

Voices from 2017

AIC provides an intimate setting for interacting with other folks in the industry, whereas other conferences you may not know anyone you meet in the hallways.
Ryan Boyd, Co-Founder, MotherDuck
Priya Nair at the panel discussion