A look back on 2016
The talks that shaped AI Council 2016.
2016 Featured Talks
Highlights from AI Council 2016 — the talks that defined the year.

Unified Pipeline Architecture: The Evolution of Data Processing at Spotify
Erin Palmer · Spotify

VC Panel - The Present Future of Data-Oriented Startups | DataEngConf NY '16
Matt Hartman, David Beyer, Evan Nisselson

To Get the Value, Ditch the Hype
Nick Ursa · The New York Times

The Future of Column-Oriented Data Processing with Arrow and Parquet
Julien Le Dem · Datadog
All 2016 Talks
Every session from 2016 — filter by topic, speaker, or company.
Unified Pipeline Architecture: The Evolution of Data Processing at Spotify
In Spotify Creator we strive to provide artists an accurate view of what their fan base is. We crunch numbers for 100 Million platform monthly active users daily and compute dashboards for millions of artists. As a result we process terabytes of data daily and condense it to about ~200 GB of data. In this talk we will discuss the evolution of our pipelines and how we made them more resilient to the irregularities of data as well as external failures.

To Get the Value, Ditch the Hype
The hype cycle around every minor innovation in data storage and data science can lead to technology and technique crushes that drift us from the critical business goals. In this talk we'll talk about concrete examples of business demands at The New York Times and how that affected our technology choices. Also, how to effectively roll out a new stack in an environment that spans 40 years of data storage systems and siloed social structure. Dirty laundry will be aired... Bubbles will be burst.

The Future of Column-Oriented Data Processing with Arrow and Parquet
In pursuit of speed and efficiency, big data processing is continuing its logical evolution toward columnar execution. Julien Le Dem offers a glimpse into the future of column-oriented data processing with Arrow and Parquet. A number of key big data technologies have or will soon have in-memory columnar capabilities. This includes Kudu, Ibis, Drill and many others. Modern CPUs will achieve higher throughput using SIMD instructions and vectorization on Apache Arrow’s columnar in-memory representation. Similarly Apache Parquet will provide storage and I/O optimized columnar data access using statistics and appropriate encodings. For interoperability, row-based encodings (CSV, Thrift, Avro) combined with general-purpose compression algorithms (GZip, LZO, Snappy) are common but inefficient. Julien explains why the Arrow and Parquet Apache projects define standard columnar representations that allow interoperability without the usual cost of serialization. This solid foundation for a shared columnar representation across the big data ecosystem promises great things for the future. Julien discusses the future of columnar data processing and the hardware trends it can take advantage of. Arrow-based interconnection between the various big data tools (SQL, UDFs, machine learning, big data frameworks, etc.) will allow using them together seamlessly and efficiently without overhead. When collocated on the same processing node, read-only shared memory and IPC avoid communication overhead; when remote, scatter-gather I/O sends the memory representation directly to the socket, avoiding serialization costs; and soon RDMA will allow exposing data remotely.

The Trials and Tribulations of Scaling Data Science and Engineering
At BuzzFeed we view data scientist and data engineers as partners working together to use our enormous data sets to understand a rapidly changing media landscape. Like any partnership there have been some bumps. In this talk we share what we've found doesn't work -- and should be avoided -- and what we've found works well. We'll also cover in detail specific strategies to make a better working relationship between data science and data engineering.

SystemML & Spark: a Framework for Scalable Data Science Algorithm Development
Apache systemML is IBM's open source project that interfaces with the Spark Context, allowing for simple expression of numerical algorithms. This is an ideal platform for Data Science, especially when there is an interest in specializing machine learning algorithms for specific challenges. The platform is extremely flexible, and enables complex numerical algorithms to be expressed in a simple and readable syntax, while preserving scalability for heavy duty computations. The parallelization details are optimized through the powerful cost based optimization engine in systemML.

Stop Obsessing about Data Infrastructure
Big Data is more popular than ever. The data world as we once knew it has changed, and that has impacted our daily lives. We spend too much time chasing the latest hype, and not enough time focusing on our data. In this talk we'll cover: how fundamentally data applications have evolved the impact of the shift from on-premise to cloud the rise of real time data Join us as we discuss, debate, and argue how the time has come to abstract away the data layer.

Statistical and Computational Challenges of Real-Time News Clustering
The proliferation of online news has been a challenge for both journalists, news consumers and policymakers who wish to take the pulse of the world because it is infeasible to manually browse and summarize this ever growing amount of data. Real time story clustering can solve this problem but it demands statistically robust and computationally efficient methodologies. As such, although heavily researched, it remains one of the open questions for both computational experts and media researchers how to group articles that cover the same news event upon their publication.

Reinforcement Learning for Data Scientists
Reinforcement Learning has received an enormous amount of attention in the machine learning community recently, with milestones such as the defeat of the world champion Go player to Google's AI making headlines, and companies like OpenAI promoting research. In this context, it makes sense to explore the role that Reinforcement Learning can play in a data scientist toolbox.

Python Data Wrangling: Preparing for the Future
In this talk, I'll discuss the current status of the pandas project and where we are planning to take it in the near future. I'll also talk about related work in data interoperability, such as Apache Arrow, designed to bring together the Python and Big Data / Hadoop worlds.

Showing 10 of 24
The voices that shaped 2016
Learn from the engineers at OpenAI, NVIDIA, and Anthropic who are moving the industry forward.

Josh Schwartz
Co-founder & CEO, Phaselab

Julien Le Dem
Principal Engineer, Datadog

Wes McKinney
Principal Architect, Posit

Alex Robinson
Core Developer of CockroachDB, Cockroach Labs

Amit Sharma
Researcher, Microsoft Research

Andy Pavlo
Assistant Professor of Databaseology, Carnegie Mellon University

Ashley Miller
Director Of Engineering, Datadog

Ben Wellington
Researcher, Two Sigma
Supported by leaders in AI infrastructure





















Voices from 2016


AIC provides an intimate setting for interacting with other folks in the industry, whereas other conferences you may not know anyone you meet in the hallways.
Ryan Boyd, Co-Founder, MotherDuck











