PyTorch Conference Europe 2026

Sponsored Session: Fault-Tolerant Training: How We Build Rel... Cyril Konkratenko & Maurits de Groot

21:15 · 07 Apr 2026 – 08 Apr 2026 · YouTube

About this talk

This talk covers fault-tolerant training for large-scale distributed AI workloads, highlighting Nebius's innovative strategies to maintain reliability amidst infrastructure failures. The speaker discusses critical reliability metrics such as goodput, mean time between failures (MTBF), and mean time to recovery (MTTR) in conjunction with automated practices like health checks, workload isolation, and state recovery. By presenting insights from production cluster results, the session illustrates how these techniques significantly reduce interruptions, improve recovery processes, and enhance the stability and efficiency of long-running AI systems.