Tutorial: DRA-matically Simple: On-Demand GPUs for MLOps - Doug Smith & Miguel Duarte Barroso
About this talk
This talk features Doug Smith and Miguel Duarte from Red Hat as they present a hands-on exploration of Kubernetes Dynamic Resource Allocation (DRA) for MLOps. With DRA now generally available in Kubernetes 1.34, the session covers how pods can request specialized hardware, such as GPUs and FPGAs, with automatic device scheduling. Attendees will engage with k8shazgpu, a DRA driver designed to enhance GPU sharing for AI/ML developers, and utilize the open-source vLLM inference framework to manage GPU allocation. The session offers insights from the user perspective, the cluster admin's viewpoint, and a developer's look at creating a custom DRA driver.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32