KubeCon + CloudNativeCon Europe

Tutorial: DRA-matically Simple: On-Demand GPUs for MLOps - Doug Smith & Miguel Duarte Barroso

1:15:02 · 23 Mar 2026 – 26 Mar 2026 · YouTube

About this talk

This talk features Doug Smith and Miguel Duarte from Red Hat as they present a hands-on exploration of Kubernetes Dynamic Resource Allocation (DRA) for MLOps. With DRA now generally available in Kubernetes 1.34, the session covers how pods can request specialized hardware, such as GPUs and FPGAs, with automatic device scheduling. Attendees will engage with k8shazgpu, a DRA driver designed to enhance GPU sharing for AI/ML developers, and utilize the open-source vLLM inference framework to manage GPU allocation. The session offers insights from the user perspective, the cluster admin's viewpoint, and a developer's look at creating a custom DRA driver.