AI Council 2016

A look back on 2016

The talks that shaped AI Council 2016.

2016 New York — Talks

All 2016 Talks

Every session from 2016 — filter by topic, speaker, or company.

· Talk

In Spotify Creator we strive to provide artists an accurate view of what their fan base is. We crunch numbers for 100 Million platform monthly active users daily and compute dashboards for millions of artists. As a result we process terabytes of data daily and condense it to about ~200 GB of data. In this talk we will discuss the evolution of our pipelines and how we made them more resilient to the irregularities of data as well as external failures.

Unified Pipeline Architecture: The Evolution of Data Processing at Spotify
· Talk

The hype cycle around every minor innovation in data storage and data science can lead to technology and technique crushes that drift us from the critical business goals. In this talk we'll talk about concrete examples of business demands at The New York Times and how that affected our technology choices. Also, how to effectively roll out a new stack in an environment that spans 40 years of data storage systems and siloed social structure. Dirty laundry will be aired... Bubbles will be burst.

To Get the Value, Ditch the Hype
· Talk

In pursuit of speed and efficiency, big data processing is continuing its logical evolution toward columnar execution. Julien Le Dem offers a glimpse into the future of column-oriented data processing with Arrow and Parquet. A number of key big data technologies have or will soon have in-memory columnar capabilities. This includes Kudu, Ibis, Drill and many others. Modern CPUs will achieve higher throughput using SIMD instructions and vectorization on Apache Arrow’s columnar in-memory representation. Similarly Apache Parquet will provide storage and I/O optimized columnar data access using statistics and appropriate encodings. For interoperability, row-based encodings (CSV, Thrift, Avro) combined with general-purpose compression algorithms (GZip, LZO, Snappy) are common but inefficient. Julien explains why the Arrow and Parquet Apache projects define standard columnar representations that allow interoperability without the usual cost of serialization. This solid foundation for a shared columnar representation across the big data ecosystem promises great things for the future. Julien discusses the future of columnar data processing and the hardware trends it can take advantage of. Arrow-based interconnection between the various big data tools (SQL, UDFs, machine learning, big data frameworks, etc.) will allow using them together seamlessly and efficiently without overhead. When collocated on the same processing node, read-only shared memory and IPC avoid communication overhead; when remote, scatter-gather I/O sends the memory representation directly to the socket, avoiding serialization costs; and soon RDMA will allow exposing data remotely.

The Future of Column-Oriented Data Processing with Arrow and Parquet
· Talk

At BuzzFeed we view data scientist and data engineers as partners working together to use our enormous data sets to understand a rapidly changing media landscape. Like any partnership there have been some bumps. In this talk we share what we've found doesn't work -- and should be avoided -- and what we've found works well. We'll also cover in detail specific strategies to make a better working relationship between data science and data engineering.

The Trials and Tribulations of Scaling Data Science and Engineering
· Talk

Apache systemML is IBM's open source project that interfaces with the Spark Context, allowing for simple expression of numerical algorithms. This is an ideal platform for Data Science, especially when there is an interest in specializing machine learning algorithms for specific challenges. The platform is extremely flexible, and enables complex numerical algorithms to be expressed in a simple and readable syntax, while preserving scalability for heavy duty computations. The parallelization details are optimized through the powerful cost based optimization engine in systemML.

SystemML & Spark: a Framework for Scalable Data Science Algorithm Development
· Talk

Big Data is more popular than ever. The data world as we once knew it has changed, and that has impacted our daily lives. We spend too much time chasing the latest hype, and not enough time focusing on our data. In this talk we'll cover: how fundamentally data applications have evolved the impact of the shift from on-premise to cloud the rise of real time data Join us as we discuss, debate, and argue how the time has come to abstract away the data layer.

Stop Obsessing about Data Infrastructure
· Talk

The proliferation of online news has been a challenge for both journalists, news consumers and policymakers who wish to take the pulse of the world because it is infeasible to manually browse and summarize this ever growing amount of data. Real time story clustering can solve this problem but it demands statistically robust and computationally efficient methodologies. As such, although heavily researched, it remains one of the open questions for both computational experts and media researchers how to group articles that cover the same news event upon their publication.

Statistical and Computational Challenges of Real-Time News Clustering
· Talk

Reinforcement Learning has received an enormous amount of attention in the machine learning community recently, with milestones such as the defeat of the world champion Go player to Google's AI making headlines, and companies like OpenAI promoting research. In this context, it makes sense to explore the role that Reinforcement Learning can play in a data scientist toolbox.

Reinforcement Learning for Data Scientists
· Talk

In this talk, I'll discuss the current status of the pandas project and where we are planning to take it in the near future. I'll also talk about related work in data interoperability, such as Apache Arrow, designed to bring together the Python and Big Data / Hadoop worlds.

Python Data Wrangling: Preparing for the Future

Showing 10 of 24

2016 New York — Speakers

The voices that shaped 2016

Learn from the engineers at OpenAI, NVIDIA, and Anthropic who are moving the industry forward.

Josh Schwartz, Co-founder & CEO, Phaselab

Co-founder & CEO, Phaselab

Julien Le Dem, Principal Engineer, Datadog

Principal Engineer, Datadog

Wes McKinney, Principal Architect, Posit

Principal Architect, Posit

Alex Robinson, Core Developer of CockroachDB, Cockroach Labs

Core Developer of CockroachDB, Cockroach Labs

Amit Sharma, Researcher, Microsoft Research

Researcher, Microsoft Research

Andy Pavlo, Assistant Professor of Databaseology, Carnegie Mellon University

Assistant Professor of Databaseology, Carnegie Mellon University

Ashley Miller, Director Of Engineering, Datadog

Director Of Engineering, Datadog

Ben Wellington, Researcher, Two Sigma

Researcher, Two Sigma

2016 New York — Sponsors

Supported by leaders in AI infrastructure

Snowflake
TextQL
HEX
Databricks
Braintrust
ClickHouse
Snorkel
Datalinks
Airbyte
Render
Turbopuffer
DigitalOcean
CockroachDB
bem
Preset
LanceDB
Chalk
Unstructured
MotherDuck
Crux
TOPK
2016 New York — Testimonials

Voices from 2016

AIC provides an intimate setting for interacting with other folks in the industry, whereas other conferences you may not know anyone you meet in the hallways.
Ryan Boyd, Co-Founder, MotherDuck
Priya Nair at the panel discussion