Keynote: Rules of the Road for Shared GPUs: AI Inference Scheduling at Wa... M. Muralikrishnan (ASL)
About this talk
This keynote features Mukund Muralikrishnan, Staff Engineer at Wayve, who discusses the complexities of AI inference scheduling in large, multi-tenant Kubernetes clusters. As the demand for AI inference workloads grows, the need for predictable access to GPUs becomes crucial, especially when managing diverse workloads that compete for GPU resources. The speaker explains the limitations of default Kubernetes scheduling and presents Kueue, a Kubernetes-native queuing and admission control solution designed to enhance scheduling efficiency in shared GPU environments. This approach not only allows for predictable GPU allocations but also improves cluster utilization and minimizes operational challenges. The session concludes with insights on how frameworks like Ray integrate into this model, supporting the scaling of Wayve's AI Driver platform.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32