Project Lightning Talk: Scheduling at the Edge of Reason: Multi-Cluster AI & OCM - August Simonelli
About this talk
This talk focuses on Open Cluster Management (OCM) and its multi-cluster capabilities, emphasizing features such as dynamic scoring and federated learning. The speaker explains how dynamic scoring automates resource management by deploying agents that gather data from clusters and communicate with a scoring API to make placement decisions dynamically. Additionally, the talk covers federated learning using the OCM hub to address high egress costs at the edge, incorporating frameworks like Flower and Open FL for deployment. The speaker also provides project updates, including alignment with SIG MC’s cluster profile API and the introduction of a queue add-on for batch learning. The advancements in the placement decision API are highlighted for their role in simplifying workload management, and the speaker expresses excitement about the project's progress and its application for incubation.
Full transcript
Goodemiddag. Hello everyone. Um, before I start I have a question. And I think it's obvious from what we've just seen in the sessions, but um, who manages more than one cluster? Right? And who manages more than one type of cluster? Yeah, okay, so it's most of you. So the main thing I want to convey here about Open Cluster Management is that it's multi-cluster and that we should
all be thinking multi-cluster. You're hearing talks from everyone about that kind of stuff, so that's good. Today we're going to talk about a few features from Open Cluster Management. That's dynamic scoring, an update on federated learning, and some project updates. Okay, dynamic scoring is a framework framework for automating resource management on the fly from your metrics. So think of it this way. We deploy agents to the
to the clusters, right? And these are deployed as an add-on. The agents then gather data from the clusters or from an external source. So you can feed this in and that is sent to a scoring API. You bring the scoring API, we give you logic on how to build it. That is going to constantly evaluate metrics based on whatever you want to use. So you I do
have a demo. You can visit us in the project pavilion for that demo. Um, essentially what that does will is create placement and placement decisions inside of Open Cluster Management. So this is happening constantly in the background. Whatever logic you want to evaluate on, it will go ahead and put that together and create this placement. What does that mean as a user? As you can see up
here, you have a workload you want to deploy. You send that to an orchestrator, which is then reading the placement. And again, the placement is constantly dynamically updating. So you might have an LLM backing it or some other fancy piece of logic, whatever you want to do. Um, and then that allows you to send the workloads down to the clusters. It's a simple concept, but it makes
it very easy to use um, and to very simply do dynamic uh All right, next thing I want to talk about is federated learning. Um what we're doing is using the OCM hub as the aggregator to help solve the high uh egress costs of doing training logic at the edge. So, we've used a flower auto add-on in the beginning here um and that helps to automate the
super node life cycle. Basically, it's it's an add-on. So, an add-on in open cluster management is a way of easily deploying software to all your managed clusters. So, that handles that. And then we abstract the orchestration into a federated learning CRD uh that reconciles the state into a manifest work. What does all that mean? Basically, it takes all your federated learning logic and turns it into open
cluster management logic, which is just uh just placement. So, it's very simple and it's very open. Um you can do it in many different ways. It's framework agnostic, so if you're not using flower, that's okay. You can use uh Open FL, whatever you want to do. Um but what's really cool, and this is something that, you know, we we love is that we're it uses the placement
API. So, we're not just dropping labels on thing things, we're using cluster claim-based scheduling. That is very helpful because it makes sure that the the um the work placement is always kept uh inside the behind the cluster boundary. So, you can do that work out on the edge without all the extra um all all the extra challenges. Okay, project update time. So, super exciting. We're aligning with
SIG MC's cluster profile API um to allow us to be a credential source, making it really easy to bring other projects into OCM. We want to be modular. We're a nice core, but we want to bring other things in so that all your multi-cluster challenges can be solved easily and openly. Speaking of add-ons, we have a queue add-on, which is really neat because it makes it easy
to install queue if you want to do batch learning or if you want to do kind of job management. But when you go to multi-queue, it can be hard to kind of build it out. You can do that with us. Super exciting add-ons are now V1 beta one, so you know what that means. Basically, they're stable, they're ready to go. We feel they're they're they've hit that
point. Finally, kept 5313 is our placement decision API. The open cluster management community has worked really hard on that. It's important because it helps decouple scheduling selection from workload dispatch, meaning it keeps a cleaner system allowing you not to have to worry about the internal logic of scheduling in something like open cluster management, but instead just throw your work at it. What else? Incubation. Super excited. We've
applied for incubation. There's not much we can do now. We just wait. It's It was a really fun process to understand how we can put that together, and we're excited to be doing that. And finally, we released our 1.0 a while back, and we're now coming forward with 1.3. So, that's the most of it. What I'm excited to also say is we are five, so that's exciting.
We have a project pavilion where I can show you a really detailed demo of dynamic scoring used with MCP. And I don't have cake, but I don't know if anyone knows what that is. Does anyone know what that is? Shout it out. That's right. Yeah, so that's it. That's all we have, Docker valve, and thanks a lot. Awesome. Thank you very much.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32