Making Topology-Aware Scheduling Practical for AI Workloads: From Discovery to Simula... Weizhou Lan
About this talk
This talk discusses making topology-aware scheduling practical for AI workloads in large-scale inference clusters. The speaker, Weizhou Lan from Daocloud, addresses challenges related to efficient GPU utilization and dynamic RDMA networking amidst heterogeneous GPU interconnect technologies. Key topics include dynamic topology discovery and health detection across multiple layers, priority-based placement for optimal communication paths, and cost-effective simulation of large, multi-level topologies using Kwok. This practical approach to topology-aware scheduling aims to enhance resource efficiency without incurring significant hardware costs.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32