Keynote: Rules of the Road for Shared GPUs: AI Inference Scheduling at Wayve - Mukund Muralikrishnan
About this talk
This keynote discusses the challenges of scheduling AI inference workloads in large, multi-tenant Kubernetes clusters, focusing on GPU access and resource management. The speaker, Mukund Muralikrishnan, a Staff Engineer at Wayve, explains how Kubernetes supports various inference tasks, highlighting the need for predictable access to GPU resources. The talk covers the limitations of default Kubernetes scheduling and details the utilization of Kueue, a Kubernetes-native solution that enhances GPU cluster management by improving allocation predictability and cluster utilization. The session concludes with insights on how frameworks like Ray integrate into their architecture as Wayve evolves its AI Driver platform.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32