Sponsored Session: Fault-Tolerant Training: How We Build Rel... Cyril Konkratenko & Maurits de Groot
About this talk
This talk covers fault-tolerant training for large-scale distributed AI workloads, highlighting Nebius's innovative strategies to maintain reliability amidst infrastructure failures. The speaker discusses critical reliability metrics such as goodput, mean time between failures (MTBF), and mean time to recovery (MTTR) in conjunction with automated practices like health checks, workload isolation, and state recovery. By presenting insights from production cluster results, the session illustrates how these techniques significantly reduce interruptions, improve recovery processes, and enhance the stability and efficiency of long-running AI systems.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17