The Token Slice: Implementing Preemptive Scheduling Via Chunked Decod... Maroon Ayoub & Kellen Swain
About this talk
In this talk, Maroon Ayoub from IBM and Kellen Swain from Google present the concept of Chunked Decoding as a solution to the challenges faced in production large language model (LLM) serving. They address the critical trade-off between maximizing throughput through continuous batching and maintaining service level agreements (SLAs) by mitigating Head-of-Line (HoL) blocking. The speakers introduce a sidecar implementation for PyTorch-based servers that allows for fine-grained preemption in a multitasking environment, effectively pausing or swapping requests without losing key-value (KV) cache. Attendees will gain insights into how varying chunk sizes can optimize priority handling and tail latency, providing a framework for advanced scheduling techniques in future model servers.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17