GPUs on Kubernetes: What Actually Happens When You Request Nvidia... Gulcan Topcu & Daniele Polencic
About this talk
This talk explores the intricacies of GPU scheduling in Kubernetes, specifically addressing what occurs when a user requests Nvidia.com/gpu: 1 in their pod specification. The speakers, Gulcan Topcu and Daniele Polencic from LearnKube, trace a GPU workload from start to finish, detailing how device plugins communicate with the scheduler and how the container runtime integrates GPU resources. Attendees will gain insights into the challenges of sharing a single GPU among multiple pods, examining solutions such as time-slicing, MIG hardware partitioning, and software enforcement. This session is designed for those curious about the technical underpinnings of Kubernetes and GPU utilization, no prior GPU knowledge is required.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32