Keynote: The Unbearable Lightness of (Agentic) Evaluations - Besmira Nushi
About this talk
This talk by Besmira Nushi, Senior Manager of AI Research at NVIDIA, explores the evolving discipline of evaluating large language models in the context of agentic tasks, which involve planning and executing autonomous actions in the real world. The speaker addresses the methodological challenges in reproducing agentic evaluations, highlighting issues such as differences in reference implementation, error handling, and tooling definitions. Additionally, the session discusses the infrastructural requirements necessary for conducting these evaluations efficiently at scale. The talk concludes with a discussion on emerging best practices aimed at enhancing consistency in measurement pipelines.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17