KubeCon + CloudNativeCon Europe

Achieving Resilient Multi-Cluster AI Inference on Kubernetes With Kar... Wei-Cheng Lai & Han-Ju Chen

28:56 · 23 Mar 2026 – 26 Mar 2026 · YouTube

About this talk

This talk focuses on achieving resilient multi-cluster AI inference on Kubernetes using Karmada and KubeRay. The speakers, Wei-Cheng Lai from Bloomberg and Han-Ju Chen from Anyscale, discuss the challenges of AI inference at scale, such as bursty traffic, uneven GPU supply, and regional latency, which a single cluster cannot adequately address. They present a practical blueprint for orchestrating Kubernetes fleets with Karmada, highlighting policy-based placement, replica spreading, and automated failover, while integrating Ray Serve-based inference managed by KubeRay. Attendees will learn when to implement a multi-cluster strategy, how to encode Karmada policies to meet service-level objectives, and how to operate Ray Serve effectively, along with receiving a reference architecture and useful templates.