KubeCon + CloudNativeCon Europe

LLM Inference at Scale: Orchestrating Prefill-Decode Disaggregation - Zhonghu Xu

32:23 · 23 Mar 2026 – 26 Mar 2026 · YouTube

About this talk

This talk covers the innovative approach to orchestrating Prefill-Decode disaggregated large language model (LLM) workloads in Kubernetes, presented by Zhonghu Xu from Huawei Technologies Co., Ltd. The speaker explains how the Prefill-Decode disaggregation architecture enhances key performance metrics like Time-To-First-Token and Time-Per-Output-Token by separating the prefill and decode stages. The session showcases Kthena's lightweight API designed for hierarchical role-based deployments, which allows for dynamic adjustments of instance ratios and improved scheduling through network topology awareness, utilizing tools like Volcano or Kueue for optimal inference performance.