Lightning Talk: KV-Cache Centric Inference: Building a State-Aware... Maroon Ayoub & Martin Hickey
About this talk
This talk focuses on KV-cache centric inference, highlighting the significance of state management in LLM inference, particularly the KV-cache which is crucial for reducing latency in production environments. The speakers, Maroon Ayoub and Martin Hickey from IBM Research, discuss how traditional optimizations around compute have become less effective as the bottleneck shifts to state management. They introduce a novel approach where cache serves as the central organizing principle of the serving platform, encompassing tiered memory management, cross-replica visibility, and cache-aware scheduling. The session details how the tools llm-d and vLLM are utilized to implement these concepts, illustrating practical applications through benchmarks, deployment patterns, and insights from their experience building a robust KV-cache centric platform.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17