AI Council 2018

A look back on 2018

The talks that shaped AI Council 2018.

2018 San Francisco — Talks

All 2018 Talks

Every session from 2018 — filter by topic, speaker, or company.

· Talk

In-Memory Data Grids (IMDGs) are the backbone of some of the most data-intensive workloads in the world. If you are booking travel, making a stock trade, or buying a home, chances are an IMDG is involved. This talk will focus on the architecture of In-Memory Data Grids by diving into the internals of Apache Geode, a popular, open-source IMDG.  Through understanding the internal architecture and characteristic of these systems we will discover the data engineering problems they solve, and when / when not to use them. We will also get hands on with Apache Geode and see how it can be used to speed up a legacy relational database.

What the heck is an In-Memory Data Grid?
· Talk

Weld is a new open source project from Stanford to accelerate data-intensive applications by as much as 100x. It does so by JIT-compiling parallel code and optimizing across functions within a single library as well as across different libraries, so developers can write modular code and still get close to bare metal performance without incurring expensive data movement costs. Weld uses a common representation to capture the structure of data-parallel workloads such as SQL, machine learning, and graph analytics and then optimizes across them using adaptive optimizer that takes into account hardware characteristics. Weld contains APIs in Python and C, and can be integrated it into a variety of widely used analytics frameworks such as Spark SQL, TensorFlow, and Pandas. Even though individual functions in these libraries are optimized, the cost of moving data across these functions can cause order of magnitude slowdowns in the whole workflow compared to a tuned implementation written in C. For example, even though TensorFlow uses highly tuned linear algebra functions for each of its operators, workflows that combine these operators can be 16x slower than hand-tuned code. Similarly, workflows that perform relational processing in Spark SQL or Pandas, numerical processing in NumPy, or a combination of these tasks spend much of their time in data movement across processing functions and could run between 2x and 300× faster if optimized end to end. We demonstrate how Weld can be incrementally integrated into these libraries by porting only the most impactful operators first without breaking compatibility with other operators in the library, and without changing the API of the libraries (so users do not need to change their application code). We also show how Weld speeds up existing workloads in these frameworks by up to 30x and enables speed-ups of two orders of magnitude in applications that combine them. The Weld library and Weld-enabled versions of the Pandas and NumPy libraries are available to download on PyPi. Weld is open source at http://weld.stanford.edu.

Weld: Accelerating Data Science by 100x
· Talk

Technical VCs get real about what it actually takes to raise money as an engineer-founder — why patents and "proprietary data" matter far less than execution, how to treat sales like a recruiting funnel, and why brutal honesty about your product's current state beats an oversold vision every time.

VC Panel Talk
· Talk

Uber’s mission is to provide transportation as reliable as running water, everywhere, and for everyone. To fulfil this mission, Uber relies heavily on making data-driven decisions at every level. Thus, we need to store more and more data as the business grows in addition to providing faster, more-reliable, and more-performant access to our analytical data. The Uber data platform is built around Hadoop ecosystem and stores more than 100 PetaBytes of data. This talk will dive into our Hadoop platform journey at Uber over the past few years, where we are standing now, and what we are building next. We started by emphasizing on data reliability, solved scalability and ease-of-use challenges and are currently focusing on faster data as well as improved efficiency.We'll look behind the scene at the current technology landscape at Uber including various big data solutions like Hadoop, Spark, Hive, Presto, Kafka, Avro, and Vertica as well as Uber's open-sourced applications and services such as Hudi, Marmaray, and Peloton. We'll dive into the technical aspect of how data freshness can be reduced from 24 hours down to minutes, ease-of-use be improved by adding a Hadoop dispersal service, GDPR regulatory requirements be addressed by providing update functionality for existing append-only columnar Hadoop data, and efficiency be improved by unifying ingestion services/pipelines. You’ll leave the talk with greater insight into how things work at Uber and will be inspired to re-envision your own data platform. In this talk we reflect on Uber’s journey with scaling our Data Infrastructure: how did we have to reinvent ourselves scaling from 1PB to 10PB to 100PB and beyond while reducing latency from 24 hours to 3h to 1h to 10 minutes, what tools did we have to make and open source to make this happen, and at what point should you think about building Data Platform.

Uber’s Data Journey: 100+PB with Minute Latency
Keynote
· Keynote

Machine learning is being deployed in a growing number of applications which demand real-time, accurate, and robust predictions under heavy serving loads. However, most machine learning frameworks and systems only address model training and not deployment. Clipper is an open-source, general-purpose model-serving system that addresses these challenges. Interposing between applications that consume predictions and the machine-learning models that produce predictions, Clipper simplifies the model deployment process by adopting a modular serving architecture and isolating models in their own containers, allowing them to be evaluated using the same runtime environment as that used during training. Clipper's modular architecture provides simple mechanisms for scaling out models to meet increased throughput demands and performing fine-grained physical resource allocation for each model. Further, by abstracting models behind a uniform serving interface, Clipper allows developers to compose many machine-learning models within a single application to support increasingly common techniques such as ensemble methods, multi-armed bandit algorithms, and prediction cascades. In this talk Joey will provide an overview of the Clipper serving system and discuss their experience transforming a research prototype into an active, open source system. He will then discuss some recent work on end-to-end cost-aware resource allocation and scheduling for multi-model applications.

The Design of Systems for Real-time Prediction Serving
· Talk

Years ago when working at Amazon on shopping cart infrastructure and the precursor to DynamoDB, my co-founder and I realized that while distributed key value stores were useful for a few use-cases, we missed many of the benefits of relational databases: transactions, joins, and the power of the lingua franca of RDBMS’s: SQL. So we challenged ourselves to modernize the traditional relational database, to take a robust open source relational database and transform it into a distributed database. This talk is about my team’s journey to create a more modern relational database. I’ll talk about the distributed systems problems we had to solve in order to scale out the Postgres open source database, in order to achieve parallelism and a concomitant increase in performance. I'll describe the architecture of the distributed query planner; how we extend traditional relational algebra operators to plan distributed queries and scale reads. I’ll also describe distributed deadlock detection, and how that enabled us to scale out transactions spanning multiple machines.

Scaling a Relational Database for the Cloud-age
· Talk

The administration of medical health plans requires policy definitions that are highly complex with legal, ethical, clinical, and financial considerations. Managing and updating these policies therefore requires significant subject matter expertise, and balancing these considerations makes it difficult to make updates that satisfy all of the constraints. This talk focuses on bringing concepts from computing and language processing such as the use of custom lexers/parsers and git-integration to streamline policy management. The policy representation and translation problem is handled using a structured natural language programming (SNLP) approach which translates from a policy language usable by a healthcare administrator into a semantic serialized object. This makes it possible to build a configuration management framework for policy management that is equivalent to “safe” policy management in mission-critical regulated industries such as developing software requirements for nuclear power systems.

Safely Streamlining Healthcare Policy Management using Ideas from Structured Natural Language Processing (SNLP)
· Talk

Structured Streaming is the next generation of distributed, streaming processing in Apache Spark. Developers can write a query written in their language of choice (Scala/Java/Python/R) using powerful high-level APIs (DataFrames / Datasets / SQL) and apply that same query to both static datasets and streaming data. In case of streaming, Spark will automatically create an incremental execution plan that automatically handles late, out-of-order data and ensures end-to-end exactly-once fault-tolerance guarantees. In this practical session, I will walk through a concrete streaming ETL example where – in less than 10 lines – you can read raw, unstructured data from Kafka data, transform it and write it out as a structured table ready for batch and ad-hoc queries on up-to-the-last-minute data. I will give a quick glimpse of advanced features like event-time based aggregations, stream-stream joins and arbitrary stateful operations.

Real-Time Data Pipelines Made Easy with Structured Streaming in Apache Spark
· Talk

As Instacart has grown from a single data scientist building linear models, to multiple teams building and maintaining dozens of bespoke models, to a more mature organization collaborating across multiple fields, we’ve learned a few things the hard way. We’re open sourcing our solutions as Lore. Common Problems Information overload makes it easy to miss newly available low hanging fruit when trying to keep up with all the machine learning packages, their features, nuances and bugs — much less implementing the latest from academia. Complexity grows because valuable models are the result of many iterative insights, making individual insights harder to maintain and communicate. Repeatability is non trivial when code, data and library dependencies change constantly in modern environments. Especially when someone else wrote the original, years ago. Glue code is often mundane and tedious to write. It’s a frequent source of bugs because there is much to write, more to maintain, and all of it has low mind-share. Performance bottlenecks are easy to hit when you’re working at high levels like python or SQL. Our goal is to make machine learning approachable for Engineers and maintainable for Data Scientists. There are a lot of great libraries like numpy, pandas, scikit, tensorflow, xgboost, etc. that work together in our daily workflow. Lore is our codification of best practices that welds the valuable bits seamlessly into production models. We're open sourcing so we can learn from the community as well.

Machine Learning from Development to Production at Instacart
· Talk

Hazard / survival modeling is often under-applied given its broad use cases. For example, churn prediction is often posed as a classification problem (did churn or not), when the time component is often given short shrift (when, if ever did the churn happen?) We hope to argue that hazard modeling is a better fit for these types of problems; spread general awareness of survival modeling, metrics, and data censoring; and describe how Opendoor uses these models to estimate our holding times for homes and mitigate risk, detailing scalability and other technical challenges we had to overcome.

Hazardous Models and Risk Mitigation in Real Estate

Showing 10 of 26

2018 San Francisco — Speakers

The voices that shaped 2018

Learn from the engineers at OpenAI, NVIDIA, and Anthropic who are moving the industry forward.

George Fraser, Co-Founder & CEO, Fivetran

Co-Founder & CEO, Fivetran

Joseph E. Gonzalez, Professor, RunLLM & UC Berkeley

Professor, RunLLM & UC Berkeley

Julian Hyde, Senior Staff Engineer, Google

Senior Staff Engineer, Google

Julien Le Dem, Principal Engineer, Datadog

Principal Engineer, Datadog

Addison Huddy, Product Manager, R&D, Pivotal

Product Manager, R&D, Pivotal

Asif Khalak, Director of Data Science, Collective Health

Director of Data Science, Collective Health

Austen Head, Senior Data Scientist, Quid

Senior Data Scientist, Quid

Chris Hartfield, Sr. Data Platform Engineer, Clover Health

Sr. Data Platform Engineer, Clover Health

2018 San Francisco — Sponsors

Supported by leaders in AI infrastructure

Snowflake
TextQL
HEX
Databricks
Braintrust
ClickHouse
Snorkel
Datalinks
Airbyte
Render
Turbopuffer
DigitalOcean
CockroachDB
bem
Preset
LanceDB
Chalk
Unstructured
MotherDuck
Crux
TOPK
2018 San Francisco — Testimonials

Voices from 2018

AIC provides an intimate setting for interacting with other folks in the industry, whereas other conferences you may not know anyone you meet in the hallways.
Ryan Boyd, Co-Founder, MotherDuck
Priya Nair at the panel discussion