A look back on 2023
The talks that shaped AI Council 2023.




















2023 Featured Talks
Highlights from AI Council 2023 — the talks that defined the year.

Writing Unit Tests for Data Science Code
Dr. Nile Wilson · Microsoft

Why People Started Testing Their Models and Data in CI / CD Pipelines
Shir Chorev · Deepchecks

What I Don’t Want To Exist In The Data World In 5 Years
Ben Rogojan · Seattle Data Guy
Using Spatial Indexes to perform geospatial analytics at massive scale
Matt Forrest · Carto
All 2023 Talks
Every session from 2023 — filter by topic, speaker, or company.
Writing Unit Tests for Data Science Code
If you type 'unit test' into your favorite search engine, you will receive a lot of information for Software Engineering, but very little guidance for Data Science code. In Data Science, the small piece of code that you want to test also needs to take in data, training a model, or evaluating a model, but all of these steps are complicated and consist of many smaller units. In this talk, Dr. Nile Wilson will share her Software Engineering best practices for testing Data Science Code and some of the common scenarios for data, like mocking calls or mocking data. This talk is for anyone bridging the gap between Software Engineering and Data Science, anyone in MLOps, or anyone productionalizing data.

Why People Started Testing Their Models and Data in CI / CD Pipelines
As machine learning models are becoming more common in production, organizations are recognizing the significance of continuous validation, and are integrating automated testing into their CI/CD pipelines to ensure that their models remain relevant and are trustworthy. However, with constantly changing data and black-box logic, testing these models can be a daunting task. In this talk, we'll explore the common pitfalls of ML models and best practices for testing them. We'll demonstrate how to use the deepchecks open source package to validate models and data during the research and CI/CD phases . By the end of the session, attendees will be equipped with the knowledge and tools to integrate ML model testing into their workflow, ensuring their models are reliable and maintainable.

What I Don’t Want To Exist In The Data World In 5 Years
Whether consulting or working as an employee there are certain tools, patterns and practices I think many of us would like to disappear in the next few years. Many of them delay projects, frustrate data engineers and yet we continue to rely on them. Whether it be transfering data via SFTP or joining teams without coding standards, some companies, even those that may be considered cutting edge, still have these patterns. Why can't we move on? In this talk I will explore some of these tools, patterns and practices as well as why I hope I don’t see them around in a few years.

Using Spatial Indexes to perform geospatial analytics at massive scale
Regularly working with big spatial datasets in your organization? Often too big to render, analyze and interpret? One way of dealing with this challenge is the use of a support geography - a simplified representation of your data that is more manageable. Spatial Indexes are multi-resolution, hierarchical grids that are “geolocated” by a short reference string, rather than a complex geometry. Smaller to store and faster to process; they out-compete conventional geometries for complex data analysis and visualization. In CARTO's recent guide (Spatial Indexes 101) you can find out more about how they work and why you should be using them in your analysis. Download here!
URGENT! Help these Pets Find Homes: Working Across Teams in DataHub
Long Tail Companions (a hypothetical pet adoption service) is in crisis– their data infrastructure has ground to a halt and they are unable to process any adoptions. In this 45 minute hands-on workshop you’ll be split into teams responsible for fixing problems in different parts of their pipeline. Each team will use DataHub to ingest data from a different platform( DBT, Looker, or Snowflake) and locate process and pipeline failures using the information supplied by other teams. Working together in DataHub, you’ll find the source of the failure, and get the adoption process moving again!
Train, Deploy, and Run a ML model using Python, Snowpark and Streamlit
In this session, we will train a Linear Regression model to predict future ROI (Return On Investment) of variable advertising spend budgets across multiple channels including search, video, social media, and email using Snowpark for Python and scikit-learn. By the end of the session, you will have an interactive web application deployed visualizing the ROI of different allocated advertising spend budgets. During this hands-on session, we will: - Set up your favorite IDE (e.g. Jupyter, Visual Studio Code) for Snowpark and ML - Analyze data and perform data engineering tasks using Snowpark DataFrames - Use open-source Python libraries from a curated Anaconda channel with near-zero maintenance or overhead - Deploy ML model training code to Snowflake using Python Stored Procedures - Create and register Python User-Defined Functions (UDFs) for inference - Create Streamlit web application that uses the UDF for real-time prediction based on user input
This App Ends Tantrums: How ML, NLP, and Five Minutes of Playtime Help Parents, Caregivers, and Children Enjoy Life Together
One in five kids has a mental or behavioral disorder, but only 15% have access to care, and the current supply of trained therapists barely covers that demand. Happypillar is a digital therapeutic app that provides evidence-proven behavioral intervention to all at scale. Learn how we combine ML, ASR, NLP, and other technologies with the expertise of our founding clinical play therapist to offer accurate and real-time personalized feedback, all with compliant security processes and the strictest privacy controls.

The state of cross-company data exchange
Data exchange is integral to every business partnership. However, data exchange practices today are mostly manual, prone to data leaks, difficult to validate, inherently impossible to monitor, and costly to audit. In this talk, we present an overview of the variety of methods enterprises use to share and transfer data. We compare these methods along the vectors of security, simplicity, and speed. We conclude by offering best practices and solutions to some of the challenges.

The things I wish I knew -- What I've gotten right and wrong from startups to the White House, and the world ahead
Thursday, March 30th @ 9:00 am It's been a decade since we've published our Harvard Business Review (HBR) article Data Scientist the Sexiest Job of the 21st Century. It's one of the 100 most downloaded articles in the history of HBR and shows how far we've come in a decade. In that decade from building companies to the White House, to leading the COVID response, I'll share my key lessons I wish I knew a decade ago.

The Story of DevRel at Snowflake: How We Got Here
Felipe Hoffa and Daniel Myers joined Snowflake as Developer Advocates in 2020 and have been integral to the company's incredible rise. In this talk, Felipe and Daniel will present an honest take of their wildly different approaches to Developer Relations and how both have been critical in building Snowflake's world-class developer community and ecosystem from the ground up. You'll learn how they define DevRel KPIs & metrics and hear about daily challenges they face and lessons learned along the way. You might even get inspired to become a Developer Advocate after understanding the different ways to engage with the Snowflake community and what's next for Snowflake Developer Relations.

Showing 10 of 85
The voices that shaped 2023
Learn from the engineers at OpenAI, NVIDIA, and Anthropic who are moving the industry forward.

David Wilson
Co-Founder & CEO, Hunch Tools

Hamel Husain
Machine Learning Consultant, Parlance Labs

Julian Hyde
Senior Staff Engineer, Google

Julien Le Dem
Principal Engineer, Datadog

Noelle Saldana
Director of Product Management, Data Science & Analytics, Data Science & Analytics

Paul Blankley
Co-founder & CTO, Zenlytic

Pedram Navid
Head of Data Engineering & DevRel, Dagster Labs

Timothy Chan
Head of Data, Statsig











