PyTorch Conference Europe 2026

Lightning Talk: Not All Tokens Are Equal: Semantic KV-Cache for Agen... Maroon Ayoub & Hyunkyun Moon

10:27 · 07 Apr 2026 – 08 Apr 2026 · YouTube

About this talk

This talk, presented by Maroon Ayoub from IBM Research and Hyunkyun Moon from moreh, explores the concept of Semantic KV-Cache in the context of serving agentic AI workloads. The speakers highlight the limitations of current inference stacks that treat the key-value cache as a flat tensor buffer, failing to account for the structural intricacies of agentic systems. They introduce a new architectural approach that categorizes cache blocks into SystemPrompt, ToolDefinition, and ReasoningBranch, allowing for differentiated caching strategies. By implementing these lifecycle-aware policies, the talk illustrates how to reduce recomputation and alleviate the challenges associated with traditional PyTorch serving stacks.