PyTorch Conference Europe 2026

Beyond the Theory: What Actually Breaks When You Scale Your Disaggregat... Ekin Karabulut & Ron Kahn

20:51 · 07 Apr 2026 – 08 Apr 2026 · YouTube

About this talk

This talk focuses on the challenges of scaling disaggregated PyTorch models for inference in high-demand environments. The speakers, Ekin Karabulut and Ron Kahn from NVIDIA, discuss the concept of disaggregated inference, which involves dividing workloads across various workers to enhance GPU utilization and performance. They address critical factors that influence scaling, such as the necessity of coordinated scaling between prefill and decode workers and the impact of network topology on KV-cache transfers. The session delves into real-world experiences from deploying vLLM and SGLang on Kubernetes, highlighting lessons learned, as well as approaches that led to performance improvements through the integration of specialized APIs like LWS and Grove with frameworks such as llm-d and Dynamo.