Project Lightning Talk: Evolving KServe: The Unified Model Inference Platform For Both... Yuan Tang
About this talk
This talk covers the evolution of KServe, a unified model inference platform that supports both predictive and generative AI. The speaker, Yuan Tang, discusses the critical need for scalable and flexible model serving infrastructure as organizations adapt to generative AI. Key topics include the transition from custom-built deployments to cloud-native, Kubernetes-based platforms, as well as challenges in productionizing large language models, such as inference efficiency and cost optimization. The session highlights the latest release of KServe, which introduces features like a new CRD designed for large language model serving, support for disaggregated inference architectures, and enhanced caching capabilities.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32