KubeCon + CloudNativeCon Europe

Optimizing LLM Inference for the Rest of Us - Abdel Sghiouar, Google

32:36 · 23 Mar 2026 – 26 Mar 2026 · YouTube

About this talk

This talk, delivered by Abdel Sghiouar from Google, focuses on optimizing Large Language Model (LLM) inference for organizations that lack the extensive resources of major tech companies. The speaker discusses the challenges posed by high costs and limited availability of GPU and TPU accelerators while highlighting the need for a balance between performance, scalability, and cost-efficiency in deploying LLMs. The session covers practical strategies for optimizing Kubernetes clusters, including container and model optimization, accelerator management, data and storage solutions, network load balancing, and observability metrics. Attendees will gain actionable insights to enhance the performance and cost-effectiveness of their AI-powered applications within Kubernetes environments.