AI Engineering

Eval Agents: How to Solve Error Cascades in Agents

Agents or RAG chatbots are multi-turn AI systems. Multi-turn means interacting back-and-forth with humans. These systems face a fundamental challenge: errors compound and cascade with each interaction. In this talk, we'll go through real-world examples of agents failing in spectacular ways when one step goes wrong - overconfidence, manipulation, looping actions, and more. After doing so, we'll examine how agent builders use "eval agents" tuned on real-world interactions to evaluate agents and even use them as verifiers to improve performance in production! By the end of the talk, you'll have learned about the new world of trajectory evaluation needed to evaluate agents accurately.