LLM Inference at Scale: Orchestrating Prefill-Decode Disaggregation - Zhonghu Xu
About this talk
This talk covers the innovative approach to orchestrating Prefill-Decode disaggregated large language model (LLM) workloads in Kubernetes, presented by Zhonghu Xu from Huawei Technologies Co., Ltd. The speaker explains how the Prefill-Decode disaggregation architecture enhances key performance metrics like Time-To-First-Token and Time-Per-Output-Token by separating the prefill and decode stages. The session showcases Kthena's lightweight API designed for hierarchical role-based deployments, which allows for dynamic adjustments of instance ratios and improved scheduling through network topology awareness, utilizing tools like Volcano or Kueue for optimal inference performance.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32