KubeCon + CloudNativeCon Europe

Optimizing Error Recovery for Cost-Ef... Radostin Stoyanov, Andrey Velichkevich & Viktória Spišáková

21:14 · 23 Mar 2026 – 26 Mar 2026 · YouTube

About this talk

This talk focuses on optimizing error recovery for cost-efficient distributed AI model training using Kubernetes, presented by Radostin Stoyanov from the University of Oxford, Andrey Velichkevich from Apple, and Viktória Spišáková from Masaryk University. The speakers address the challenges of achieving scalable and fault-tolerant distributed AI training, particularly with interactive GPU workloads like Jupyter notebooks. They introduce transparent GPU checkpointing as a solution integrated with Kubernetes to enhance cluster utilization and improve cost efficiency. The session also examines how checkpoint policies work with Kubernetes APIs such as Kueue, JobSet, and TrainJob, enabling users to take advantage of preemptible spot instances for their AI model training needs.