PyTorch Conference Europe 2026

From Responses To Trajectories: Multi-Turn and Multi-Environ... Kashif Rasul & Sergio Paniego Blanco

23:25 · 07 Apr 2026 – 08 Apr 2026 · YouTube

About this talk

This talk covers the advancements in post-training large language models (LLMs) using reinforcement learning, specifically focusing on trajectory-based optimization rather than static prompt-response pairs. The speakers, Kashif Rasul and Sergio Paniego Blanco from Hugging Face, discuss multi-turn and multi-environment Generalized Reward Preference Optimization (GRPO) training, which allows LLMs to gain interactive agent-like experiences through simulated environments and multi-step reasoning tasks. They also explore how the TRL framework, designed for PyTorch, supports scalable workflows and integrates with simulated environments like OpenEnv. Attendees will learn about key design patterns, trajectory batching, and advantage computation, which enhance alignment, reasoning, and generalization in LLMs for agentic applications.