PyTorch Conference Europe 2026

Lightning Talk: Trinity Large - Torchtitan on 2000+ B300s - Matej Sirovatka, Prime Intellect

10:34 · 07 Apr 2026 – 08 Apr 2026 · YouTube

About this talk

This talk covers the use of torchtitan for scaling the training of ultra-sparse mixture-of-experts models across over 2,000 GPUs. The speaker, Matej Sirovatka from Prime Intellect, walks through the pre-training of Trinity Large, a 400 billion parameter mixture-of-experts model, with an emphasis on maximizing throughput and reducing the impact of hardware failures. Key challenges such as fault tolerance, large-scale distributed training, and ensuring determinism are discussed, along with solutions implemented using torchtitan. The session concludes with insights and common pitfalls to avoid during large-scale training initiatives.