KubeCon + CloudNativeCon Europe

Tutorial: KV-Cache Wins You Can Feel: Building AI-Aware... Tyler S, Kay Y, Vita B, Nili G & Maroon A

1:21:22 · 23 Mar 2026 – 26 Mar 2026 · YouTube

About this talk

This hands-on tutorial covers the construction of AI-aware LLM routing on Kubernetes, presented by Tyler Michael Smith from Red Hat, along with Kay Yan from DaoCloud and team members from IBM. Participants will learn how to deploy a distributed vLLM cluster and benchmark its performance while visualizing the inefficiencies of cache-blind routing. The session demonstrates how to replace the default Service with the Kubernetes Gateway API and implement llm-d, a framework optimized for distributed LLM inference. Attendees will walk away with practical lab setups, dashboards, and an understanding of integrating cache-aware routing into production AI systems to enhance performance and reduce costs.