Project Lightning Talk: Next-Gen AI Orchestration With Volcano On Kubernetes - Zhonghu Xu
About this talk
This talk covers next-generation AI orchestration using Volcano on Kubernetes. The speaker provides an overview of the Volcano project, originally developed for batch processing, which has evolved to support heterogeneous device management and various types of AI workloads. With the rise of large language models (LLMs) and their unique resource requirements, Volcano aims to unify scheduling for AI training and inference. The technical roadmap for Volcano includes updates such as a new agent scheduler for handling latency-sensitive workloads and enhanced network topology awareness for precise pod scheduling. The introduction of the HyperNode component allows for better modeling of network architecture, while sub-projects like Casino enhance native serving capabilities for AI inference frameworks and agent workloads.
Full transcript
Uh, I'm from Huawei Technology and this is currently one of the Volcano maintainer. Um, so let's begin. Uh, today my topic is next-gen AI orchestration with Volcano on Kubernetes. So let's get a quick overview of Volcano. Some listeners might have uh, listened uh, already be familiar with the Volcano before but others may not. So let's quickly get a quick look look at the Volcano project. Uh, Volcano
was originally from the query batch and has has been worked as a cloud not cloud native batch system. Uh, it can perform uh, queue management and gang scheduling and and do has a lot of rich for ecosystem and support uh, a lot of heterogeneous uh, device. A few years ago most clusters uh, were dominated by the batch scale uh, training but in 2023 more and more uh,
large language models are emerging uh, LLM training uh, is experienced uh, explosive growth. So now we run also into the uh, disaggregated uh, LLM inferences uh, where prefer and decode have different resource requirement and currently we also have the agent workloads which is uh, bursty short-lived. So different workload characters uh, pose challenging to the scheduler uh, and Volcano has been uh, widely used in AI training uh,
since its incubation. So Volcano will continue to evolve uh, and to support uh, LLM inference and agent workloads to build as a unified scheduling platform. So here the Volcano technical uh, roadmap for this year and uh, beyond. From the So the front level of volcano will uh, not only support the VC job the no native will will also support more native workloads to post support massive AI
training AI inference agent workloads and collaborate with other more communities such as reinforcement learning areas inference framework and agent agent frameworks. So and at and at the bottom volcano will support more heterogeneous devices as a unified heterogeneous device allocated pool and support more features in the network architecture topology aware scheduling to improve the AI workloads performances. So let's follow the roadmap to introduce what we updated in
in the last year. One of them is that volcano is originated from the batch scheduler can do the batch workloads but the its framework is just a schedule batch workloads per second. So the but may be sensitive latency sensitive and I need to burst it. So we build build a new scheduler called agent scheduler in the previous version coordinated with the existing batch concurrently scheduling in the
same control plane and and we build a shutting controller to manage the node CRD to coordinate the dynamic the node pools. And what's news in the volcano network topology aware scheduling? Let's quickly review the volcano network topology aware scheduling, which is some some of you may have already used before. What we saw is that large scale AI lack unified topology abstraction requiring multi multi master intra and
then internal configuration for maximum performances. And we build the HyperNode CRD to model the switch or the same performance network domain to to to model it to allow the scheduler know the downside network architecture and to schedule pods more precisely to improve the AI training influence And the HyperNode to create it manually maybe complicated, so we also provide a HyperNode pluggable auto disk discovery mechanism. We We
have already supported Nvidia fab unified fabric manual discover and the existing node node labels discovery. And we also will extend to more discovery in the futures. And and currently most of scheduler use the node labels to represent the network topology and combine their own CRD to specify the topology. But we wish to promote the HyperNode as the Kubernetes first class citizen because the HyperNode has its status
can do can show the switch congested status the bandwidth of the switch the status of the in the HyperNode. So we we wish to extend the HyperNode as a citizen because the currently we don't have the entity to specify or model the switch or the topology domain. But in the long run we wish to we wish to expand the happy new as the Kubernetes first class citizen
and we we call for co-creation for the happy new initiated. And we also have built two sub project. One is called Casino. It's a native serving platform. It support VM. You've gone a little long there, buddy. You've got to you're over a minute over so we can wrap this one up. Okay. Hurry up. Go. Go. Yeah, let's go let's quickly uh Casino is a EM serving platform.
It supports lots of inference frameworks and it can do auto scaling on the token sprout token sprout and support a lot of scaling mechanism and it it has intelligent root router can do KB cache preview cache PD group aware routing and it's collaborate with the volcano scheduling can do again scheduling and schedule topology aware scheduling. And another sub project is called the agent cube. It can do
AI agent workload It's built on the six sandbox today but we will introduce agent primitives and border integrations work with more framework. And agent cube also provides out of box SDK to allow you directly reviews warm tool build the AI agent workloads. Thank you. All right, thank you.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32