Optimizing Reinforcement Learning at Trillion-Parameter Scale - Songlin Jiang
About this talk
This talk covers the implementation and optimization of reinforcement learning at the trillion-parameter scale using Mixture-of-Experts reasoning models. The speaker discusses system design strategies for large-scale RL training systems, including the use of LoRA for expert layers, sharding and fusion under tensor, pipeline, and expert parallelism, and the parameter synchronization process between Megatron and vLLM. The session also addresses training-inference mismatch in MoE RL, explaining the limitations of traditional mitigation methods and detailing the fixed Router Replay R3 approach developed to align routing decisions across vLLM, veRL, and Megatron.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17