From Responses To Trajectories: Multi-Turn and Multi-Environ... Kashif Rasul & Sergio Paniego Blanco
About this talk
This talk covers the advancements in post-training large language models (LLMs) using reinforcement learning, specifically focusing on trajectory-based optimization rather than static prompt-response pairs. The speakers, Kashif Rasul and Sergio Paniego Blanco from Hugging Face, discuss multi-turn and multi-environment Generalized Reward Preference Optimization (GRPO) training, which allows LLMs to gain interactive agent-like experiences through simulated environments and multi-step reasoning tasks. They also explore how the TRL framework, designed for PyTorch, supports scalable workflows and integrates with simulated environments like OpenEnv. Attendees will learn about key design patterns, trajectory batching, and advantage computation, which enhance alignment, reasoning, and generalization in LLMs for agentic applications.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17