It Works in the Demo. Will It Work in Production? Evaluating and Debugging AI Agent - Apurva Misra
About this talk
This talk addresses the challenges of ensuring an AI agent's reliability in production environments, where real users, varying prompts, and unstable tools reveal vulnerabilities. The speaker demonstrates how to transform a functional agent into a trusted system by defining reliability concepts for multi-step, tool-using behaviors and exploring success, partial success, and failure modes. The session emphasizes building production-relevant evaluation suites through the creation of golden datasets and scenario-based action path tests, while also covering the structuring of evaluation runs to identify brittleness and silent failures. Attendees will learn to implement scoring methods, regression testing, and monitoring systems for ongoing assessment as changes occur in prompts, tools, or APIs.
More from this event
See all 126 talks →
AI Is Not the Risk. Architectural Drift Is - Sunil Kalkunte
17:39
Breaking the Monolith: Tesco’s Journey to Federated GraphQL with xAPI - Vishwas Chandrashekar
29:13
A Practical Introduction to LangChain4j - Venkat Subramaniam
1:01:28
Beyond the AI Models: How Lowe’s is Building the Store That Knows - Swaroop Shivaram
13:59