Tutorial: KV-Cache Wins You Can Feel: Building AI-Aware... Tyler S, Kay Y, Vita B, Nili G & Maroon A
About this talk
This hands-on tutorial covers the construction of AI-aware LLM routing on Kubernetes, presented by Tyler Michael Smith from Red Hat, along with Kay Yan from DaoCloud and team members from IBM. Participants will learn how to deploy a distributed vLLM cluster and benchmark its performance while visualizing the inefficiencies of cache-blind routing. The session demonstrates how to replace the default Service with the Kubernetes Gateway API and implement llm-d, a framework optimized for distributed LLM inference. Attendees will walk away with practical lab setups, dashboards, and an understanding of integrating cache-aware routing into production AI systems to enhance performance and reduce costs.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32