Ethan Rosenthal

Member of Technical Staff, Runway

Ethan Rosenthal is a Member of Technical Staff at Runway, an applied AI research company focused on multimedia content creation, where he builds engineering systems to accelerate the work of research scientists. His career spans diverse roles across AI, machine learning, and data science - from training language models at Square to developing recommendation systems at seed-stage ecommerce startups. Before working in tech, Ethan was an actual scientist and got his PhD in experimental physics from Columbia University.

Ethan Rosenthal

Sessions / 2025 / 1 talk

  • While it is often easier to simply throw more data at a problem, scale is not all you need when building multimodal foundation models. Data quality continues to be just as important as data quantity, and supporting “data-centric AI” requires lowering the barrier to data curation as much as possible. However, multimodal data curation presents unique requirements compared to conventional machine learning or business intelligence data management systems. The data is heterogeneous, ranging from scalars to embedding arrays to entire compressed videos. While the dataset sizes in terms of number of rows are not quite Big Data™, the number of bytes is massive with high columnar variance. Given the storage size, it’s infeasible to construct and copy new training datasets for each model training job; training jobs must query the core datasets without copying them. Finally, large scale distributed training jobs require fast random access which bumps up against limitations of typical solutions like partitioned parquet files. In this talk, I will discuss how we built a petabyte-scale, multimodal feature lakehouse. This lakehouse supports analytical querying as well as serving features for large scale distributed training jobs, such as those that were used for training Runway’s recent foundation models like Gen-3 Alpha.

Sessions / 2018 / 1 talk

  • Machine learning has revolutionized the capability of businesses to create personalized experiences via real-time, individual predictions and recommendations. But what happens when one must make thousands of decisions for thousands of individuals at the same time? At Dia&Co, a plus-size women’s styling service, we recently faced such an obstacle when building out a brand new product line for the business. This talk will explore how we combined modern machine learning with classical operations research techniques to scale personalization in the face of constraints inherent to a retail business. The basics of operations research will be introduced before demonstrating how to solve a simple version of our real-world problem using all open source libraries. I will then reveal the gory details of productionizing this work, from testing to gracefully handling failures of convergence. Finally, I will cover the journey from the coldest of starts, with zero data, to synthesizing machine learning with the operations research problem.

Ready to take this stage?

The next edition is programmed by practitioners. Tell us what you built and what you learned.

Apply to be a speaker