Achieving Resilient Multi-Cluster AI Inference on Kubernetes With Kar... Wei-Cheng Lai & Han-Ju Chen
About this talk
This talk focuses on achieving resilient multi-cluster AI inference on Kubernetes using Karmada and KubeRay. The speakers, Wei-Cheng Lai from Bloomberg and Han-Ju Chen from Anyscale, discuss the challenges of AI inference at scale, such as bursty traffic, uneven GPU supply, and regional latency, which a single cluster cannot adequately address. They present a practical blueprint for orchestrating Kubernetes fleets with Karmada, highlighting policy-based placement, replica spreading, and automated failover, while integrating Ray Serve-based inference managed by KubeRay. Attendees will learn when to implement a multi-cluster strategy, how to encode Karmada policies to meet service-level objectives, and how to operate Ray Serve effectively, along with receiving a reference architecture and useful templates.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32