Building a Flexible Data Platform for LLM Training Data
World-class LLMs are trained on trillions of tokens, and data quality is critical in determining the ultimate performance of these systems. At Cohere, we train LLMs and retrieval (RAG) systems from scratch using datasets created through complex ingestion, preprocessing, and distillation pipelines. This talk will cover what we know about how data drives LLM performance, as well as the data platform we use to manage hundreds of datasets, from ultra-niche finetuning datasets to petabyte-scale web data, and automate the measurement and enforcement of data quality at scale. Here’s what practitioners will walk away from in this talk: What the science and our own experiences show us about data for LLMs A detailed understanding of the unique architecture we built to manage datasets for training complex NLP systems The anatomy of an LLM training data pipeline, and how we build data quality evaluation into these pipelines Practical lessons for how they can implement best practices for their LLM training infrastructure