Lightning Talk: Not All Tokens Are Equal: Semantic KV-Cache for Agen... Maroon Ayoub & Hyunkyun Moon
About this talk
This talk, presented by Maroon Ayoub from IBM Research and Hyunkyun Moon from moreh, explores the concept of Semantic KV-Cache in the context of serving agentic AI workloads. The speakers highlight the limitations of current inference stacks that treat the key-value cache as a flat tensor buffer, failing to account for the structural intricacies of agentic systems. They introduce a new architectural approach that categorizes cache blocks into SystemPrompt, ToolDefinition, and ReasoningBranch, allowing for differentiated caching strategies. By implementing these lifecycle-aware policies, the talk illustrates how to reduce recomputation and alleviate the challenges associated with traditional PyTorch serving stacks.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17