Make Too Much Knowledge Just Enough. Massive Scale RAG and GraphRAG with Open Source
RAG systems that work in the real world are not just the trivial extract, vector search, and rerank systems that the simplistic "Introductions to RAG" suggest. After this talk, you will understand how to think about the design and construction of real world RAG and GraphRAG systems that can scale to hundreds of millions of documents or billions of vectors. You will learn about the complex orchestration of multiple libraries. You will also learn how to use tools and frameworks that use open standards like OpenTelemetry or OpenInference to help you monitor and debug these complex RAG orchestrations. Topics will include discussions of scalable RAG/GraphRAG architectures, complex extraction flows, embedding model and re-ranking considerations. We will dive deep into integration between various libraries like Ray, LangChain, LlamaIndex, DSPy, Phoenix, Weaviate, PgVector, GraphRAG, LangGraph, AirFlow, KFP and vLLM to form a cohesive solutions that actually scale. We will discuss the patterns and anti-patterns Cake has learned building and deploying these systems for real customers. If time permits, we will address advanced topics like complex table-detection/extraction for financial data, complex agentic flows to handle heterogeneous datasets, etc.