Redefining SLIs for LLM Inference: Managing Hybrid Cloud wit... Christopher Nuland & Hilliary Lipsig
About this talk
This talk focuses on redefining Service Level Indicators (SLIs) for Large Language Model (LLM) inference and managing hybrid cloud environments through the use of vLLM and LLM-D. The speakers, Christopher Nuland and Hilliary Lipsig from Red Hat, address the challenges SREs face as traditional performance metrics fall short in the context of LLM applications. They discuss the importance of metrics such as Time-to-First-Token, cache hit ratio, and GPU utilization, and explain how vLLM and LLM-D can enhance observability and scalability for LLM inference. Attendees will learn to establish new SLIs, instrument distributed inference systems using tools like Prometheus, OpenTelemetry, and Grafana, and effectively integrate LLM telemetry into Kubernetes SRE practices.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32