Beyond the Theory: What Actually Breaks When You Scale Your Disaggregat... Ekin Karabulut & Ron Kahn
About this talk
This talk focuses on the challenges of scaling disaggregated PyTorch models for inference in high-demand environments. The speakers, Ekin Karabulut and Ron Kahn from NVIDIA, discuss the concept of disaggregated inference, which involves dividing workloads across various workers to enhance GPU utilization and performance. They address critical factors that influence scaling, such as the necessity of coordinated scaling between prefill and decode workers and the impact of network topology on KV-cache transfers. The session delves into real-world experiences from deploying vLLM and SGLang on Kubernetes, highlighting lessons learned, as well as approaches that led to performance improvements through the integration of specialized APIs like LWS and Grove with frameworks such as llm-d and Dynamo.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17