Route, Serve, Adapt, Repeat: Adaptive Routing for AI Inference Workl... Nir Rozenbaum & Kellen Swain
About this talk
This talk covers adaptive routing strategies for AI inference workloads in Kubernetes, presented by Nir Rozenbaum from Red Hat and Kellen Swain from Google. The speakers discuss the limitations of traditional static routing methods, such as traffic splitting and session stickiness, which fail to account for the dynamic nature of inference requests and evolving cluster conditions. They introduce the K8s Gateway API Inference Extension, highlighting how it utilizes real-time data like queue length and cache utilization to optimize routing. Attendees will gain insights into the impact of these adaptive strategies on latency, efficiency, and cost-effectiveness in managing Kubernetes inference workloads.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32