Keynote: Rules of the Road for Shared GPUs: AI Inference Scheduling at Wa... M. Muralikrishnan (ASL)
About this talk
This talk discusses the use of AI for autonomous driving at Waabi, focusing on the challenges of processing vast amounts of driving data. The speaker highlights the operation of Kubernetes clusters across Azure, which support over 100,000 concurrent workloads and utilize more than 1,000 GPU nodes. The recent adoption of Kueue, a Kubernetes-native job queuing system, is presented as a solution for improving resource management and ensuring fairness in a multi-tenant environment. The implementation of Kueue has significantly increased GPU cluster utilization from 85% to 97%, while also reducing average wait times for smaller teams by nearly 70%. The talk emphasizes the importance of open source contributions and concludes with an invitation to learn more about Kueue at KubeCon.
Full transcript
Hello everyone. It's great to be here this morning. AI inferencing seems to be the theme for the day, but I'm going to be talking about that as well, but just not about text. At Waabi, we build end-to-end AI for autonomous driving. One AI model that can learn from vast amounts of data and generalize across vehicles and regions. Our fleet of vehicles, including from our partners in dashcam
OEM and taxis, collect thousands of hours of driving data every day. And to process that requires massive compute. We operate Kubernetes clusters across multiple Azure regions, each running more than 1,000 GPU nodes. These clusters process a huge variety of workloads across different teams, GPU hardware, priorities, and SLAs. At peak, we we handle about 100,000 concurrent workloads. We rely heavily on open source to operate at this scale.
I want to thank all of the contributors here and the broader open source community for building the technologies that make this possible. Today, I want to highlight our recent adoption of Kueue to schedule our AI inferencing workloads. It has helped us improve the utilization of one of our most expensive and scarce GPU clusters from 85% to over 95%. Kubernetes is incredible, but we lack the granular controls
to ensure fairness in a highly competitive multi-tenant environment. During periods of high churn, tens of thousands of pending pods caused degradation in performance in the kube scheduler. That's where Kueue came in. Kueue is a Kubernetes-native job queuing system with advanced controls for resource management. It complements kube scheduler rather than trying to replace it. It's highly scalable. We have scale tested it up to 100,000 pods, and it
required no code changes to integrate with our existing workloads. Each team gets a guaranteed allocation, and when they don't use it, other teams can burst into that capacity, and Kueue ensures it's distributed in a fair manner. Soon after the launch, we saw average wait times drop across all the teams. Especially for the smaller teams which used to get starved, the wait times came down by almost 70%.
Even during heavy bursts, Kubernetes scheduler stayed highly performant because it only sees the pods that can actually be allocated on the nodes. As a result, the utilization of our GPU cluster jumped from 85% to 97%. But the best part is that we went from desiring to implement Kueue and having it running at full production scale in less than a month, including all alerting and monitoring. If Kueue
sounds interesting, I would like you to I would recommend you to check out the talks by Kueue contributors here at KubeCon later today and throughout the week. And if what we do at Wave sounds exciting, we are hiring globally. I'll be at the Microsoft booth later today. Please come talk to me. Thank you and have a great KubeCon everyone.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32