PyTorch Conference Europe 2026

Optimizing Reinforcement Learning at Trillion-Parameter Scale - Songlin Jiang

24:34 · 07 Apr 2026 – 08 Apr 2026 · YouTube

About this talk

This talk covers the implementation and optimization of reinforcement learning at the trillion-parameter scale using Mixture-of-Experts reasoning models. The speaker discusses system design strategies for large-scale RL training systems, including the use of LoRA for expert layers, sharding and fusion under tensor, pipeline, and expert parallelism, and the parameter synchronization process between Megatron and vLLM. The session also addresses training-inference mismatch in MoE RL, explaining the limitations of traditional mitigation methods and detailing the fixed Router Replay R3 approach developed to align routing decisions across vLLM, veRL, and Megatron.