KubeCon + CloudNativeCon Europe

Making Topology-Aware Scheduling Practical for AI Workloads: From Discovery to Simula... Weizhou Lan

24:13 · 23 Mar 2026 – 26 Mar 2026 · YouTube

About this talk

This talk discusses making topology-aware scheduling practical for AI workloads in large-scale inference clusters. The speaker, Weizhou Lan from Daocloud, addresses challenges related to efficient GPU utilization and dynamic RDMA networking amidst heterogeneous GPU interconnect technologies. Key topics include dynamic topology discovery and health detection across multiple layers, priority-based placement for optimal communication paths, and cost-effective simulation of large, multi-level topologies using Kwok. This practical approach to topology-aware scheduling aims to enhance resource efficiency without incurring significant hardware costs.