Cloud Native Theater | Istio Day: Running State of the Art Inference... Jackie Maertens and Nili Guy
About this talk
This talk covers the challenges of efficiently running large language model workloads, highlighting the issues posed by limited GPU memory and shifting traffic patterns. The speakers, Jackie Maertens from Microsoft and Nili Guy from IBM, introduce LLM-D, a distributed inference system designed to address these problems using the Gateway API Inference Extension. They explain several innovative techniques such as cache-aware request routing, prefill/decode split execution, locality-aware scheduling, and dynamic worker scoring, all aimed at enhancing throughput and reducing latency for real-world inference traffic. Attendees will learn how to leverage these advancements with their existing Istio installations to optimize inference serving.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32