Joseph E. Gonzalez

Professor, RunLLM & UC Berkeley

Joseph Gonzalez is a Computer Science Professor at UC Berkeley, the co-director of the RISE and Sky Computing Labs, and a member of the Berkeley AI Research (BAIR) group. He leads the Large Models Systems (LM-Sys) and Gorilla projects which have significantly advanced open research in large language models and their supporting systems. Joseph is also co-founder and Head of AI at RunLLM.com. His research addresses problems in data systems, large language models, scalable machine learning, computer vision, robotics, autonomous driving, and graph analytics. Prior to joining Berkeley, Gonzalez co-founded Turi Inc (formerly GraphLab) based on his thesis work. Gonzalez’s innovative work has earned him significant recognition, including the Okawa Research Grant, the NSF Expedition Award, the VLDB Test-of-Time award, and NSF Early CAREER Award.

Joseph E. Gonzalez

Sessions / 2025 / 2 talks

  • This keynote panel features Denis Yarats and Joseph Gonzalez -- two pioneers bridging academic theory and practical application. Joseph Gonzalez has transformed his Berkeley research into tangible solutions through LM-Sys and Gorilla projects, now bringing his expertise in machine learning and robotics to RunLLM.com after successfully launching Turi based on his doctoral work. Denis Yarats complements this approach with his reinvention of information discovery at Perplexity, where he's leveraging his NYU PhD and Facebook AI experience to develop Comet, a revolutionary "browser for agentic search." Together, they exemplify how rigorous academic foundations can be transformed into technologies that solve real-world problems and reshape our digital interactions.

  • The Future of AGI: Building Compound AI Systems | Explore a paradigm shift in AGI development through the lens of compound AI systems that integrate multiple LLMs with specialized tools. Learn how orchestrating diverse AI components can achieve human-level performance across broad task domains, demonstrated through RunLLM's AI support engineer implementation. Features practical approaches to building general-purpose AI workflows that combine speed, accuracy, and adaptability. Includes real-world case studies showing how compound AI systems are transforming customer support and service automation.

Sessions / 2024 / 1 talk

  • Join DJ Patil, the former US Chief Data Scientist, and Joey Gonzalez, founding member of RISElab at UC Berkeley and an AI research and entrepreneurship expert, for an insightful dialogue on the multifaceted world of AI. Key discussion points include: • Regulatory Challenges in AI: A deep-dive into how the AI industry navigates its unique regulatory landscape, underlining the criticality of overseeing AI applications rather than the root technology. • The AI Startup Journey: Take a behind-the-scenes look at Joey's AI startup, Run LLM. Discover the evolution from cloud technologies to specialized developer assistants harnessing the power of advanced AI. • AI Market Trends: Stay abreast of the latest shifts in the AI market, learn about changing investor perceptions and the hurdles in demonstrating tangible value in AI applications. • Innovative AI Research: Spotlight on Joey's research group's breakthrough projects, presenting exciting advances in model delivery optimization and agentic language models. Peek into the intricate and engaging world of AI, its regulatory context, the highs and lows of AI entrepreneurship and the cusp of AI research innovation.

Sessions / 2018 / 1 talk

  • Machine learning is being deployed in a growing number of applications which demand real-time, accurate, and robust predictions under heavy serving loads. However, most machine learning frameworks and systems only address model training and not deployment. Clipper is an open-source, general-purpose model-serving system that addresses these challenges. Interposing between applications that consume predictions and the machine-learning models that produce predictions, Clipper simplifies the model deployment process by adopting a modular serving architecture and isolating models in their own containers, allowing them to be evaluated using the same runtime environment as that used during training. Clipper's modular architecture provides simple mechanisms for scaling out models to meet increased throughput demands and performing fine-grained physical resource allocation for each model. Further, by abstracting models behind a uniform serving interface, Clipper allows developers to compose many machine-learning models within a single application to support increasingly common techniques such as ensemble methods, multi-armed bandit algorithms, and prediction cascades. In this talk Joey will provide an overview of the Clipper serving system and discuss their experience transforming a research prototype into an active, open source system. He will then discuss some recent work on end-to-end cost-aware resource allocation and scheduling for multi-model applications.

Ready to take this stage?

The next edition is programmed by practitioners. Tell us what you built and what you learned.

Apply to be a speaker