Lightning Talk: Trinity Large - Torchtitan on 2000+ B300s - Matej Sirovatka, Prime Intellect
About this talk
This talk covers the use of torchtitan for scaling the training of ultra-sparse mixture-of-experts models across over 2,000 GPUs. The speaker, Matej Sirovatka from Prime Intellect, walks through the pre-training of Trinity Large, a 400 billion parameter mixture-of-experts model, with an emphasis on maximizing throughput and reducing the impact of hardware failures. Key challenges such as fault tolerance, large-scale distributed training, and ensuring determinism are discussed, along with solutions implemented using torchtitan. The session concludes with insights and common pitfalls to avoid during large-scale training initiatives.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17