Project Lightning Talk: From Idle to Ideal: Cross‑Cluster GPU Sharing with CoHDI - Takao Indoh
About this talk
This talk covers the challenges of maximizing GPU utilization across multiple Kubernetes clusters in light of increasing AI workloads and the high costs of GPUs and memory. The speaker presents a solution called Codi, which stands for composable hardware in disaggregated infrastructure. This approach allows for a shared pool of GPUs and memory that can be dynamically allocated and detached based on real-time needs. During periods of heavy inference traffic, more GPUs are allocated to inference clusters, while at night, resources are shifted to support training and batch jobs. By utilizing Codi, organizations can effectively reduce idle GPU time and lower infrastructure costs, making their hardware usage more flexible and efficient.
Full transcript
Hello everyone. Uh thank you for joining. My name is Takao Indo from Fujitsu. Today I will talk about uh from idle to ideal cross-cluster GPU sharing with Cordy. AI workload keep increasing. At the same time GPUs and memory are getting more expensive and scarcer. Many organization cannot can no longer buy dedicated GPUs for every team or every cluster. So, the key question is uh what if we
could maximize device utilization across multiple Kubernetes cluster? On the left in the picture, some clusters have idle GPUs and other clusters are GPU starved. So, we want to balance this. However there is a challenge. Traditionally, Kubernetes manages hardware only inside each cluster boundary. There is no straightforward and safe way to lend devices across the cluster. idle capacity in cluster A cannot easily help cluster B. This is
the problem we want to solve. This is an idea to resolve this problem. Here is a how it works day-to-day. We keep a shared GPU and memory pool. During the day, inference traffic is heavy, so we attach more GPUs to inference clusters. At night, training and batch jobs dominate. So, we detach the inference Uh I detach devices from inference cluster and attach it to training So, that
means attach when needed and detach when not needed. This is This reduces idle time and lower infrastructure cost and scale effect efficiently. The main idea is simple. GPU are allocated to where we they create the most value at the right time. Our approach is Codi. That means composable hardware in disaggregated infrastructure. Composable infrastructure let us create software-defined bare metal systems. With Codi, GPU can can be dynamically
attached and detach to Kubernetes nodes. We use a resource pool of hardware, GPUs, memory, and other devices and a management layer to compose the right bare metal at the right time. In short, put devices in shared pool and compose node you need. Attach or detach GPUs on on demand. This makes the hardware flexible and flexible, not fixed. Codi works with Kubernetes scheduler and DRA. We are currently
tracing tracking changes in version 1.36 and as beta. We have just released code E version 0.0.1.1. We have provided composable infrastructure emulator. You can try this. If you want to talk more, please visit us at the project pavilion. And you can also scan the QR code on the screen and to see the code and docs and join the project. So. That is thank you for listening. Thank
you. Thank you very much.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32