Intelligent Routing for Optimized Inference - Antonio Berben, Solo.io & Felipe Vicens, Telefonica
About this talk
This talk features Antonio Berben from Solo.io and Felipe Vicens from Telefonica, focusing on intelligent routing for optimized inference in cloud native environments. The speakers analyze the inefficient use of expensive GPU cards for simple queries that could be better managed by CPUs, such as batch processing and small context requests. They present a three-layer architecture that incorporates CNCF projects like Istio for secure communication, LLM-D as the inference framework, and kagent as the agent orchestrator. This architecture enables real-time analysis of request complexity, allowing for the efficient routing of simple queries to Small Language Models with under 8 billion parameters.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32