AI Council 2022

A look back on 2022

The talks that shaped AI Council 2022.

Attendees
Attendees
Sponsor, Mode
Women Speaker
Data Dinner, Pete Soderling
nick schrock
Attendees
Community Party
Speaker
Sponsor
Keynote Panel
2022 Austin — Talks

All 2022 Talks

Every session from 2022 — filter by topic, speaker, or company.

· Talk

Ryan Blue, co-creator of the Apache Iceberg project will try to convince you not to care about Iceberg: if you’re thinking about your table format, then it isn’t doing a good enough job. This session will show how Iceberg solves real-world problems that used to take hours or days of time from data engineers and analysts: Safe schema changes — no more zombie data columns Layout evolution — update table partitioning without rewriting any queries Hidden partitioning — safe and fast queries without being a DBA Future work — current frustrations and how we’re making them disappear

Why You Shouldn’t Care About Iceberg
· Talk

In this talk, Hung will reveal how Y42, an all-in-one data pipeline tool, has leveraged Git as a noSQL database to foster unparalleled collaboration opportunities between data engineers and data analysts. Y42’s decision to abandon all classical databases within their platform in favor of using Git has brought a series of major innovations — but this cutting-edge approach hasn’t been without challenges. This talk will highlight the extraordinary benefits of using Git-as-a-database to power an all-in-one data workspace, such as: Data-pipeline-as-code (including dashboard-as-code and integrations-as-code) Browser-based, high-performance implementation of Git using WebAssembly, meaning the Y42-data-pipeline can be implemented using no-code, low-code and/or code Simple end-to-end templating Easy version control + rollback + environment of job status, settings and data warehouse tables Coherent pipeline automations with one orchestration layer And it doesn’t end there… However, deciding to use this unconventional approach meant they were in for a bumpy ride. The Y42 team found themselves facing a series of challenges, including: Performance issues to save thousands of jobs inside Git Access control issues using folder paths as ids Getting the web app to seamlessly integrate Git and code The Y42 team invites you to attend this talk and learn more about using Git as a noSQL database for a new data revolution. It’s a bold statement, but they believe this talk has the potential to significantly shape the data industry for the years to come.

Using GIT as a NoSQL Database for Fine-Grained Control Over the Data Pipeline
· Talk

Datasets are one of the fundamental concepts in data work: as data practitioners, we use the word all the time in colloquial day-to-day conversations. Many different data tools have independently converged on similar concepts. However—like many concepts in the modern data stack—the exact meaning, properties, and capabilities of Datasets differ in subtle but important ways. Concepts like Datasets will define the next generation of data work. The cornerstone of “the modern data stack” is a new set of tools, abstractions, and metadata that map more tightly to the real work that data practitioners need to do. This panel brings together leading tool builders and practitioners in the data community to discuss that evolution. We’ll start by comparing and contrasting different approaches to Datasets. From there, we’ll branch out into an open discussion about relative strengths and weaknesses of different approaches, and alignment (or lack thereof) between tools and systems. This talk will be useful for data practitioners looking to understand how the field is evolving, and how new tools are enabling those changes.

What is a Dataset? Emerging Core Concepts in the Modern Data Stack
· Talk

Pandas has become one of the de-facto libraries for data manipulation of tabular data in the Python ecosystem. In recent years, several projects have emerged, such as Dask, Modin, and Koalas, whose goal is to reproduce the Pandas API in order to ease the learning curve for scaling data processing logic. Coupled with ML orchestration tools like Flyte, machine learning practitioners can benefit from reproducibility and data lineage tracking while using the data processing tools they are familiar with. However, as powerful as dataframes are, they can often be difficult to reason about in terms of their data types and statistical properties as data is reshaped from its raw form into one that’s ready for modeling. In this session, data science and machine learning practitioners will learn how to combine Flyte’s (LF AI & Data incubating project) rich type system and flexible DAG composition syntax with Pandera’s intuitive schema-declaration API so they can spend less time worrying about the correctness of their dataframes and more time obtaining insights and training models. This talk will first introduce Pandera (OSS project), a package that provides an expressive data validation API, and then dive into a practical case study to illustrate the benefits of integrating Pandera with Flyte.

Type-Safe Data Processing and Machine Learning Pipelines with Flyte and Pandera
· Talk

When it comes to sensitive or high-stakes use cases, such as emergency response dispatch or in air traffic management, humans’ expertise, common sense, and sensibility can’t be replaced by AI; instead, the influence of these human qualities should be augmented. When the stakes and complexity level are high, AI assistance can help humans make sense of the sheer amount of data and not be overwhelmed. In return, humans can help AI understand the larger context and be trusted to make complex decisions. In this talk, it will be explained how creating shared experiences between humans and AIs addresses the limitations of traditional AI techniques in order to create efficient human-AI teams. The speaker will present what he learned by building such intelligence ecosystems using Cogment, an open-source framework, and show the audience how they can apply these principles to their use cases.

Towards Human-AI Teaming: Intelligence Ecosystems to Tackle High-Stakes Use Cases
· Talk

What was the data analytics industry like before the rise of the data clouds? What will it look like 5, 10 years from now, and how do we get there? e.g. Are you worried about your cloud bills growing exponentially? Are your SQL users waiting on slow queries, unless they start 1000 cloud machines (because umm... "elastic computing")? Come and join us in this talk, and let's change the world together (again).

The Sky's the Limit -- RE: The Next 10 Years of Data Infra on Clouds
· Talk

Data infrastructure as a category is booming with the global market expected to exceed $100B by 2027 reflecting 200% growth since 2020. We’re also seeing an explosion of data itself, with total data produced expected to top 175 Zettabytes in the next 4 years; that’s up 350% from 2020 numbers. As a community, we’re also seeing a proliferation of data infrastructure options across both the traditional analytics stack as well as in innovative ML tooling and platforms. Data is hot, and data tooling of all kinds is even hotter. But how are we to make sense of these new data tools and new categories of tools? How should we be thinking about the modern data stack when it comes to cloud & hosted options vs. OSS and self-hosted infra? Which tools and categories of tools play together nicely, and where might we start to see consolidation across the stack? For the Next Big Opportunities in Data Infrastructure panel, we’re joined by 4 top investors who look at data infrastructure tools every day. They’ll share with us their mental models of the space and insights into how they make investment decisions across the expanding data infrastructure landscape. Founders and practitioners alike will benefit from their insights on which categories are here to stay and which are yet to be proven. Join us for a lively discussion on where the future opportunities of data infra are and where these experts think the field is headed next.

The Next Big Opportunities in Data Infrastructure
· Talk

Fifteen years ago, OLAP cubes were a critical part of every analytics and BI stack. In a time when databases were slow and compute was expensive, cubes provided an elegant solution for standardizing multi-dimensional reporting. Over the last decade, however, they’ve fallen out of favor. As warehouses have gotten bigger, faster, and cheaper, cubes no seem longer necessary. Analysis and reporting is now done directly on top of raw data, no predefined or pre-aggregated cubes required. Or are they? OLAP cubes are reappearing in the modern data stack—just in a different form and under a different name. Instead of being separate data marts built for reporting and BI, cubes are now synthetic, generalized, and on-demand. In this talk, I’ll walk through the history of OLAP cubes and their modern echoes. And I’ll explain why this is actually a good thing—and why we should actually be excited about the return of the OLAP cube.

The Return of the OLAP Cube
· Talk

Metaflow was originally developed at Netflix to provide a user-friendly platform for a wide range of ML use cases from computer vision and NLP to classical statistics. Today, Metaflow is used by hundreds of companies from real estate and finance to biotech and drones. In this talk, we give a technical overview of Metaflow, showing how it helps data scientists to develop, deploy, and operate ML projects. We walk through an exciting array of new features in Metaflow, such as support for Kubernetes, model scorecards, a new monitoring GUI, and many others.

The Modern Stack for ML Infrastructure
· Talk

In building data-centric products, companies have many architectural and cultural questions to consider. What is the right technology to use? What data design and use patterns do we want to encourage? In ZipRecruiter’s case, these questions have been answered in many different ways by many different teams since the company was founded in 2010. A decade later, we’re knee-deep in the process of bringing our data landscape under control. This talk discusses the process the company has gone through on its path towards a more mature data governance model and data-driven culture -- the challenges we’ve faced, key strategies we’ve employed, and the opportunities it’s unlocked.

The Life-Changing Magic of Data Governance

Showing 10 of 52

2022 Austin — Speakers

The voices that shaped 2022

Learn from the engineers at OpenAI, NVIDIA, and Anthropic who are moving the industry forward.

Benn Stancil, Founder, Mode

Founder, Mode

Chang She, Co-founder & CEO, LanceDB

Co-founder & CEO, LanceDB

Julien Le Dem, Principal Engineer, Datadog

Principal Engineer, Datadog

Ryan Blue, Creator of Apache Iceberg, Member of Technical Staff, Databricks

Creator of Apache Iceberg, Member of Technical Staff, Databricks

Ville Tuulos, Co-founder & CEO, Outerbounds

Co-founder & CEO, Outerbounds

Caitlin Colgrove, Founder & CTO, Hex

Founder & CTO, Hex

Abe Gong, Co-Founder & CEO, Great Expectations

Co-Founder & CEO, Great Expectations

Adi Polak, Vice President of Developer Experience, Treeverse

Vice President of Developer Experience, Treeverse