Optimizing Error Recovery for Cost-Ef... Radostin Stoyanov, Andrey Velichkevich & Viktória Spišáková
About this talk
This talk focuses on optimizing error recovery for cost-efficient distributed AI model training using Kubernetes, presented by Radostin Stoyanov from the University of Oxford, Andrey Velichkevich from Apple, and Viktória Spišáková from Masaryk University. The speakers address the challenges of achieving scalable and fault-tolerant distributed AI training, particularly with interactive GPU workloads like Jupyter notebooks. They introduce transparent GPU checkpointing as a solution integrated with Kubernetes to enhance cluster utilization and improve cost efficiency. The session also examines how checkpoint policies work with Kubernetes APIs such as Kueue, JobSet, and TrainJob, enabling users to take advantage of preemptible spot instances for their AI model training needs.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32