Optimizing LLM Inference for the Rest of Us - Abdel Sghiouar, Google
About this talk
This talk, delivered by Abdel Sghiouar from Google, focuses on optimizing Large Language Model (LLM) inference for organizations that lack the extensive resources of major tech companies. The speaker discusses the challenges posed by high costs and limited availability of GPU and TPU accelerators while highlighting the need for a balance between performance, scalability, and cost-efficiency in deploying LLMs. The session covers practical strategies for optimizing Kubernetes clusters, including container and model optimization, accelerator management, data and storage solutions, network load balancing, and observability metrics. Attendees will gain actionable insights to enhance the performance and cost-effectiveness of their AI-powered applications within Kubernetes environments.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32