About this talk
This talk introduces Hami, a GPU orchestration solution designed to optimize GPU resource usage by slicing GPUs into virtual devices. The speaker discusses Hami's capability to run multiple workloads concurrently while reducing operational costs and time to market. Noteworthy case studies, such as Baidu and SF Express, demonstrate significant cost reductions and stability improvements through Hami's dynamic memory isolation features. The session also highlights Hami's support for various accelerator types and its integration with Kubernetes, making it transparent for users. Furthermore, the talk includes a demonstration of Hami's core functionalities, showcasing the allocation of resources in real-time and emphasizing its extensibility.
Full transcript
Hey everyone, this is Reza, solution engineer at Hami. Hi everyone, I'm Dimotion and I'm the co-founder and the CTO of the Dynamia and also the maintainer of Hami project. So Hami is a GPU orchestration solution that allows you to slice GPUs into virtual devices and therefore allowing you to run more workloads in parallel and also reducing cost and time to market. So simple use case, for example,
you have a PyTorch notebook user, some support chat, both of them requesting two GPUs, one of them with 10G, the other one with 20G each. Before Hami, you would basically fill up a four-node GPU cluster. After Hami, Hami packs them nicely together and you get two extra GPUs for free. Um We have five case studies on the CNCF website right now. One of them, Baidu, actually their
inference tasks were starving out LLM and training jobs and after using Hami, they they reduced their cost by 2/3. Another one is SF Express that actually needed custom memory isolation to have their more stable. With that, I'm going to switch to the demo. Uh sorry. what you see is basically three GPU three GPU clusters, um two A100s each. One of them is configured for MIG and I'm
going to run two extra LLM workloads. on the MIG node and cannot see actually. One on the MIG node and one uh on the on the normal one. So what's going to happen is you see I'm allocating 25 gigs for the normal one, I'm allocating 10 gigs for the MIG node, which is going to do the MIG partitions automatically itself here. And then up here you see
the where you have one NVIDIA SMI output 25G. And if I reload this, you'll see that these are now scheduled here. And with that, Motion. Thanks Reza. Since there are theoretical some time for the VM module to warm up, so I will give a brief introduction to the core capabilities of Hami. So I think the most two most intriguing things about Hami is the first we have
a heterogeneous system management. It means that besides NVIDIA GPU, we also support Huawei Ascend NPU, Cambricon MLU and AWS Neuron devices and as well as Meta X, Innova her and many other sorts of accelerators. They all can be partitioned by the Hami project. And the another is how we handle the hard resource isolation inside the container. We can do this in two ways. The first one is
as you have you have seen in the demo, we have the dynamic MIG partitioning. The other is we have a self-implemented CUDA access layer, which can access the device memory allocation calls from the CUDA runtime to CUDA driver and we can do the counting here and we will reject anyone who is not illegally, yeah. And also we since we are Kubernetes native, we we we have completely
transparent to the user task and we have implemented in the DAI way as well, yeah. So with that, the LLM is still initializing. So if we look at the logs here, the LLM is almost done, so we should get an output in a second. One moment. There you go. And then of course you will see that now that it's loaded, it's no longer showing zero because the
memory is loaded and with that, let's go take a look at the ecosystem. So we have over 3,000 stars right now. We have over 500 contributors in 17 countries. Of course we have seamless integration with all the CNCF projects. Motion already mentioned all the devices that we support. We're going to support a lot more this year and then go ahead. Thanks Reza. If you are interested in
our Hami project, we have a pavilion this afternoon. Be sure to check us out and have a good cool fun. Thank you. >> Thank you. >> [applause] >> All right, all right. Platform engineering folks, where you at? Come on. Help me welcome to the stage Ben and Patrick from Backstage. Y'all having a good time talking about platform engineering? Yeah? Good week? Feeling good? All right. All right.
Hey everyone, I'm Patrick. Let's switch to the demo straight away. Yeah, it's cool. Here we go, cool. So this is a demo of Backstage's role in the age of AI. Backstage is a framework for building internal developer portals and it's a highly flexible solution that empowers your platform teams to deliver value to the rest of your organization. In this story, we're talking of an imaginary music company
and we're building a new service. We have a minimal Backstage instance set up and what I see here is a list of all the services that our team owns. Now, let's say I'm searching for an API that I want to use. I could use the search functionality and also just browse all of the APIs in my organization. Now, this recommendations API seems to be what I want.
So here on this page, I can see the API definition, information like who owns this API, what service implements it and so on, documentation as well. Now, a central part of Backstage is extensibility through plugins. And this view, which we call the entity is one of the most powerful surfaces to extend. Having customized pages for services, resources and other entities in your catalog is very powerful, especially
when it comes to day-to-day operations. You might be wondering though, how relevant is this in the AI era? Well, at Spotify, we're collecting usage metrics for both AI and our internal Backstage UI and we have seen an increase in UI usage over the last few months. And an interesting trend is the more an engineer uses AI tools in their daily workflows, the more they also use the
Backstage UI, indicating AI does not replace the UI. However, engineers is only one type of consumer of Backstage's rich catalog of and agents is of course another one. Now, I'm going to pass it over to Ben to show you more of what we have built to bring Backstage into the AI era. Hello everyone. Um I apologize in advance, it's day three of the conference and my voice
is slightly dying, but I'm going to try and get through this. Okay, um so my name is Ben and I'm going to walk through a quick of how we're integrating reusable plugin actions across multiple surfaces and meeting engineers where they're going to be spending most of their time in this new age of AI. So first off, I'm going to show you how these actions are integrated into
a new CLI and now we're enabling engineers to use these use this CLI to create skill scripts and even powering CI workflows. So I'm going to dive over to the terminal here. I'm going to first off log in to my Backstage instance, which I'm hoping is going to work. We like live demos, right? Everything's going to be fine. I get a pop-up in my Backstage instance to
approve this session, so I'm going to authorize that and that's going to give me access. Oops, ignore my desktop background, it's really not sorted. Yes, ignore that. I'm going to first off list these actions, so we can see that I have three plugins installed right now, which is the catalog, the off plugin and the scaffolder and the actions that [snorts] they provide. Now, keep in mind this
list of actions for a second. I'm going to jump into Claude real quick and I'm going to run the same thing here with the MCP and I can see I've got these same list of actions available as an MCP server. So I can get Claude to work, right? So I'm going to head back to Claude and I'm going to give it a simple prompt saying I want
to implement a view in my application where I can see the users' recommended tracks. Now, I haven't told Claude any more context than that. Like I've just given it the MCP server and it's going to go and first off search the catalog or our entire service catalog to work out if there's anything in there that smells a little bit like a recommendations API. We can see that
it's queried the catalog, it's found our recommendations API and now it's putting towards putting together a plan to be able to build this view for us. Um let's say I don't want to go and build it straight away though. I just want to go and get some more information about the API like who owns this. Like maybe I want to go and ask that team some questions
first. So I can see it's going back to the catalog, we've found the discovery team, our discovery support and now I've got the team members I can reach out to. I think that's the end of the demo. Look, can we head back to the slides real quick? Um so we know that surfaces across the web UI, the CLI and AI tools are all essential to delivering an
outstanding developer experience within your organization. We're excited about what's ahead and you can learn more about what we at Spotify are doing for Backstage at backstage.spotify.com. >> All right, two for two. Great job. All right, this last one I discovered on the airplane on the way over and I needed some fantastic dashboards and thanks to my assistants Claude and Gemini, they came up and they were spectacular.
So, um Augustin here is going to talk about Perses. Oh, it's working. Yes. Awesome, thank you. Hello everyone. Um like Jabe said, I'm Augustin. I'm working at Amadeus as a principal engineer. And it's truly amazing to be here because when I started the project 5 years ago, uh I was not thinking at all to be on stage in front of all of you to talk about this
project. So, now the big question today is what is Perses? So, Perses is a observability visualization tool. It's about displaying any kind of signals, so metrics, logs, um traces, profiling. Uh it's a CNCF standard project. I think I forgot to mention it. Um it's actively developed by Amadeus, Red Hat, and SAP. It's founded today by the European Union since last year. So, really happy that the Europe
is backing us. Um but Perses is not about just dashboarding UI. It's also providing specification, open specification, that I hope one day will be shared by any observability vendor because if it's the case, then we will have just to define one uh dashboard format and so what means any kind of any community will have to define only once their you will be able to install them in
any vendor observability you have. And that will be a really great because otherwise you need to migrate from one vendor to another vendor, you have to restart all your dashboard again and again and again. And Perses is not just about it's not just about Yeah, I hope it Yeah. It's not just about UI and specification you because sometimes you don't want to install another dashboarding UI, you
want to display the signals in your own UI, which is possible with Perses because we are providing components, React components to be able to do that. So, basically if Perses is able to display any kind of signals, now it's the case for you as well because you can embed the components in your own UI. And now let's start Let's show you how it looks like. Back to
this the demo. Oh, yeah. This is the home page of Perses where you can uh see the list of the dashboards and projects you have access for. And for this for this little demo, uh I really have a a simple setup. I have installed Minikube. I have uh on Minikube I have installed the Prometheus operator, so you have the really the bare minimum to monitor Kubernetes with
Prometheus, kube-state-metrics, node exporter. Um and also for the demo I have installed Tempo and Loki from Grafana Labs to display logs and traces. Um so, on the home page you can have access like I said to the project and for the if I'm clicking on that, I have a list of project uh a list of dashboards, sorry. Uh and as you can see, we have a bunch
of Kubernetes dashboards and that's the opportunity to show you that today with Perses you can monitor your Kubernetes cluster because we are supporting thanks to the amazing community of Perses the Kubernetes official dashboards and they are today hosted on the community mixings, but we are planning to move these dashboards to the community to the Kubernetes community. But same for Prometheus because we are supporting um the Kubernetes
dashboards, the Prometheus dashboard, Thanos, Tempo, Alert Manager, Istio, All of these dashboards are implementing in Go long using the Perses Go SDK, uh which allow us to write actually now a dashboard in Go long and not in JSONet or anywhere the templating in JSON. Um and it's also using PromQL query uh the PromQL builder, sorry. So, now you can write PromQL query in Go long using type
sets, which means if you are not writing um a good PromQL query, it won't compile. So, on the left you have the Prom the Go long definition for the API server for overview of Kubernetes. And on the on the right you have the definition in YAML. So, just let me show you because I'm running out of time. This is the API server of Kubernetes. Uh I have
also installed the Prometheus Yeah, it gives you the Prometheus metrics overview. >> [clears throat] >> And then I have the logs also installed. Yeah, awesome. It's working. And then just really quickly just to show you that we are able to display the traces as well. And actually that's it because I'm running out of time. So, if we are can go back to the slides. Maybe. Maybe not.
Okay. It's okay. Uh yeah, thank you. Uh so, that's it for the demo. If you're interested by the project, we have uh a booth at noon. Uh and uh if you want to check the project, uh the QR code will guide you to the official website perses.dev. You can follow us on social media, contact uh Slack. And that's it for me. Thank you very much. >> Three
for three. Told you it was easy. All right, really quick. If you're a CNCF project maintainer or other open source maintainer, please stand up. Give these folks a hand. >> All right, thanks everybody.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32