Chang She

Co-founder & CEO, LanceDB

Chang She is CEO/Co-founder at Eto Labs building modern data infrastructure for AI. Previously he architected the ML and experimentation stack at TubiTV as VP of Engineering. In the mythical pre-pandemic epoch, Chang was the 2nd major contributor to Pandas, CTO/Co-founder of DataPad, and a recovering financial quant.

Chang She

Sessions / 2026 / 1 talk

  • Most AI problems are really data problems. AI workloads bring with them ever larger amounts of data from multiple modalities (e.g., text, images, audio, video, sensor data). If you were indexing say, the internet, you need to solve a number of new data infra challenges: Storing large blobs and avoiding copying them over and over during processing Dealing with much larger table sizes: trillion rows with a capital T Supporting workloads like Search, Curation, and Training directly from your dataset instead of having to move data to/from point-solution systems Dealing with *really* distributed pipelines: what happens when your storage, CPUs, and GPUs are with different clouds / vendors? In this talk we will dive into detail on why it's challenging to manage trillion scale wide tables with multimodal data. We'll see why existing data infra doesn't support new these data types, workloads, or scale. And we'll do a quick under the hood peek at how Lance format and LanceDB solves these problems at a foundational level. Zooming out, we'll cover how LanceDB fits into the existing data stack alongside Iceberg. Finally, we'll talk through our roadmap and show you the big improvements we're working on in 2026. Whether you're looking to do large scale search or building the next frontier model, this will help you scale easier, get to production faster, and save on infra cost.

Sessions / 2024 / 1 talk

  • As exemplified by the success of GPT-4 and Midjourney, it is without a doubt that the future of AI is multi-modal. However, managing embeddings, images, audio, video, and other unstructured data types brings a host of new data management challenges that cannot be met with the current generation of the data stack. New data types are generally much larger in size and it is common for AI datasets to have deeply nested schemas. Workloads are also very different from traditional analytical tasks; from training, evals, debugging, and EDA, core data infrastructure needs to be natively optimized for both scans and random access, on both scalar columns and large binary or tensor columns. In this talk, we'll cover a few use cases across training, data management, and EDA of image datasets to illustrate how the current generation of data infrastructure fails for AI. We'll introduce Lance columnar format, the critical open source foundation for a new generation of data infrastructure for AI, and how we can build a new kind of Lakehouse optimized for multi-modal AI.

Sessions / 2022 / 1 talk

  • Data infrastructure for AI is broken. Repeated conversions, disjoint workflows, and manual QA of labels are just some of the pain-points experienced by any AI team trying to bring models to production. To solve these challenges, we've created a new open-source data format called Rikai. Rikai is designed specifically for AI teams to avoid the need to convert data between ETL, analysis, and training phases of the AI workflow. Using semantic-typing, Rikai also supports deep understanding of AI datasets using plain vanilla SQL. Finally, Rikai already works with existing analytics systems so it fits right into the modern data stack without the need for new compute frameworks or engines.

Ready to take this stage?

The next edition is programmed by practitioners. Tell us what you built and what you learned.

Apply to be a speaker