Lightning Talk: Inside VLLM's KV Offloading Connector: Async Memory Transfers for... Nicolò Lucchesi
About this talk
This talk covers the KV Offloading Connector in vLLM 0.11.0, which addresses the challenge of recomputing the KV-cache state for large language model (LLM) requests. The speaker, Nicolò Lucchesi from Red Hat, explains how this asynchronous, pluggable API expands KV cache capabilities by utilizing CPU DRAM to mitigate GPU memory limitations. He details the connector architecture, memory transfer tradeoffs, and the redesign of memory layout that significantly enhances offloading performance. The session provides practical guidance for enabling CPU offloading in production environments, showcasing improvements in transfer throughput and effective block sizes.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17