The Science and Practice of Open and Scalable LLM Evaluations - Grzegorz Chlebus, NVIDIA
About this talk
This talk by Grzegorz Chlebus from NVIDIA delves into the science and practice of open and scalable evaluations for large language models (LLMs). It addresses the challenges of model evaluation in the rapidly evolving AI landscape, emphasizing the importance of rigorous quality assurance and effective testing methods across multiple benchmarks. The speaker presents best practices, practical patterns, and anti-patterns for standardizing evaluation methods, while showcasing Nemo-Evaluator, an open-source tool designed for scalable evaluation. The session highlights how this tool facilitates transparent and reproducible measurement, ultimately contributing to a more open model development process for the Nemotron model family.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17