KubeCon + CloudNativeCon Europe

Intelligent Routing for Optimized Inference - Antonio Berben, Solo.io & Felipe Vicens, Telefonica

33:37 · 23 Mar 2026 – 26 Mar 2026 · YouTube

About this talk

This talk features Antonio Berben from Solo.io and Felipe Vicens from Telefonica, focusing on intelligent routing for optimized inference in cloud native environments. The speakers analyze the inefficient use of expensive GPU cards for simple queries that could be better managed by CPUs, such as batch processing and small context requests. They present a three-layer architecture that incorporates CNCF projects like Istio for secure communication, LLM-D as the inference framework, and kagent as the agent orchestrator. This architecture enables real-time analysis of request complexity, allowing for the efficient routing of simple queries to Small Language Models with under 8 billion parameters.