KubeCon + CloudNativeCon Europe

Collisions in the Dark: Illuminating the 95% of Kubeflow You Can't... Amine Lahouel & Laura Llinares

24:44 · 23 Mar 2026 – 26 Mar 2026 · YouTube

About this talk

In this talk, the speakers discuss the observability stack they are implementing at CERN to enhance the data acquisition and analysis process for high-energy physics experiments. They explain the significance of their trigger system, which filters millions of collision data points to focus on rare events for research purposes. A key component of their platform is Kubeflow, utilized for machine learning workloads, alongside various monitoring tools including Prometheus and Grafana for data visualization. The presentation highlights the importance of energy monitoring and sustainability, detailing methods for tracking energy consumption and determining the carbon footprint associated with extensive computation workloads. The speakers also share challenges faced in their observability journey and propose a collaborative approach to streamline the deployment of their observability manifest for the community.

Full transcript

Uh hello and uh welcome everybody for our talk uh collision in the dark. I forgot I'm eliminating the 95% of Kuflow that you cannot see and this will talk will explain a little bit our observability stock that we are using. Oh, I Oh, pop I didn't. So, first uh myself, I will present myself. I am I mean Lel software engineer at CERN and I have nearly uh

one decade of experience >> and you saw it already but I'm Lara. I'm 25 years old and I'm a DevOps engineer at CERN as well. >> So, do you hear me? I think yes. So before we go into the technical detail we wanted to present a bit briefly what is CERN and what is the context of all the project. So Cers CERN is the biggest um high

energy physics laboratory. It's famous for the discovery of the of the bosen and this discovery was enabled because of the large collider that we have below the swans and the French and Swiss border. This collider has as you can see on the picture many experiment that are collecting a huge amount of collision data. So those collision happens up to 40 million times per second. So this is

a lot of data. So because there's a lot of data we have to and we only care about rare and interesting event to discover new physics. Uh they use something called the trigger system. So what is it? It's um what you see here on the slide. We have the level one trigger trigger that is hardware based and we filter this 40 millions uh collision data per second

and we'll uh filter around 99% of it and all this data will then go to the level two triggers which is software based and same this second layer will also filter around 99% of the data. Then this data is sent to tape for storage and all the scientific they do their analysis on it. Uh so then we are now in what it's called run free. Uh and

we will move to something called I luminosity LHC which mean luminosity means more data. So if you want even more data you need more better trigger system and that's how the project NGT which is next generation triggers came into the picture. So it's a research and de development initiative. Uh and it's focusing on this data acquisition and this trigger system and a real-time event processing and in

our team especially we are working on the scientific computing infrastructure of NGT. So we design procure deploy and operate the hardware and software platforms that we will be supporting. Um yes so one of the software platform we have a scientific computing platform that we build uh for machine learning and it's based on cubeflow and some other promising open source component. uh we provide access to a shared

pool of resources of GPU and compute units and this can be interactive but you can also uh it's used also for job submission and we have we have a focus on machine learning workload and yesterday I put actually I put a reference to a cubecon talk that we had yesterday one of our colleagues uh they were talking about the this scientific computing platform and if you want

to see more you can see the recording probably and also there is a link to the reference architecture of NGT and I will let you continue. >> Thank you lawyer. Uh so this is analytics depiction of our user users using our platform. So when uh we deployed the platform and started on boarding users suffices they like to do uh different stuff. So uh all of them they

have their own workload and their habits and we had multiple questions from them. Some of them will be uh wondering why their training is so slow. Others will be wondering how much energy uh their workloads will be consuming and some of them will always ask for more GPU. You always need more GPU. But uh as an MLF stop the Kubernetes default stack is a little bit um

limited should I say as a default as you can see here we can uh have node mics pics cluster metric and uh some uh default dashboards that you have for the cluster node and port levels. But our users they were wondering and asking other questions. So uh they will be like why my training is so slow we don't know we don't know for that we need some

GPU metrics and operation team will be asking how much all those GPUs will consume power they will be a little bit uh worry about uh the that will go above the rack limits and stuff like this so for this we give them power and energy metics and some energy attribution also for our teams And then comes the management and the management will be asking some questions about

how efficient are we using. But what they are really saying is are we spending the money where it needs to be and for this we have some metrics about the capacity and about also the idle session that we have. Uh so I will explain a little bit the different um categories of the metrics that we have. So the platform uh usage uh we want to have active

users, the number of session notebooks and blah blah blah. And then for the GPU telemetry, all the frequency, the energy consumption, the C memory and utilization and then uh lawyer can maybe check. Yes, thank you. Because I cannot look up there. So for the resource availability also uh we need some uh real-time view of um the resources that are available in the cluster and this is very

useful for the users for the power consumption. We want to know uh how much is the cluster is consuming and related to that is the sustainability and because sun has some sustainability pledges it was really important for us to have some CO2 emission uh accountability and the final thing is the idle uh resource monitoring like I said for better efficiency. So how we did all of this?

So we came up with uh this uh very beautiful architecture diagram. Uh as you can see here on the bottom my right you have your nodes so your physical machines and on top of them we'll have different exporters for the different hardware that we have there. So we have for the Nvidia the AMD exporters and we will have some APMI and Kepler for the energy and for

the other stuff we'll have the n your typical node exporter and cube exporter and on the other side you can see we have the our kubernetes workflows on top of that you have cubeflow exporting some metrics about the jobs life cycle the notebooks the controllers and everything in between and all of is you have uh Prometheus to escape all of them and uh we are using the

cook Prometheus stack to have Prometheus deployed but also have gapana for the visualization and we'll have some dashboards there and then sending the metics about the idle sessions for the alert manager and the alert manager will notify the users by email or by notification in mattermost. So uh I will go through a little bit uh about all the metrics that we are collecting. So uh I don't

know if you know but if you go to any documentation of those exporters you'll be faced with hundreds of metrics and you'll be what metrics do I need for my users and uh our cluster. So we handpicked a little bit of those metrics for the different ters we have. So for the node exporters we have some metrics about the nodes themselves their specs uh the pod info

and pod start time for the CI advisor for the container uh running u metics and for the node exporters we have also the hardware memory and temperatures and the CPUs and then going on for the GPU exporters it's the same story for the MDN uh Nvidia one so they will also give you a lot of metrics and we need to choose ones that are interesting for our

users and then can um handle their questions and then I will give back the mic to explain a little bit uh our energy monitoring stack. >> Yes. Uh so one thing that we have to consider in the our monitoring stack is the energy and the efficiency of our hardware because we are doing physics. So any small percentage of gain in efficiency means the possible discovery of new

physics. So we emphasize our work on energy monitoring for our clusters. And so we use different tool for energy probing. We have for example IPMI. So IP IPMI is um node level metric. So it's a standard that you use for power monitoring of the hardware and it will give us um the metrics at the of the physical components. So the energy that is consuming on the other

end because uh we wanted to have also the energy accounting per at the Kubernetes granularity level. So we wanted it by pod by workload that we are running. So we use a tool that is called Kepler. It's also a CNCF project. uh I will go a bit quickly about how does Kepler does the magic. So it use uh rapper kernel modules to and sensor to collect data

and um um about CPU package about cores about the memory subsystem and so on and it also collect different uh software information about the running processes and then it does all the m magic because when you have the data about the processes it's uh easy to go back and go to the container level then the pod level and so on. So yeah, does the magic and it

gives us the energy attribution of the our workloads that we're running and then it exports all the mismetic and we can scribe them with Prometheus. also as he was mentioning I mean SASM sustainability p pledge that we want to uh that we committed to and to do that we wanted to cross our data about energy monitoring with uh the carbon data and it's really easy at CERN

because um everything that we run is on the CERN data center which is running on the French grid. So you just have to fetch an API that is public and you fetch the carbon intensity. The carbon intensity is the gram of CO2 equivalent per kilowatt hour. So if you have the also your power consumption, you just have to multiply it and you have your CO2 equivalent. And

that's how we also obtain a really good dashboard about carbon emission and I mean we >> and yeah so as you may expect with all those exporters and all those metrics we are collecting uh this was really hard on our Prometheus deployment and we had to learn that the hard way. uh we cached our cluster uh our Prometheus uh multiple times and we even lost the data

and we learned some hard lessons along the way and uh here you can see a sample of our configuration that we have now. So uh before that uh on the top one you can see a little bit uh we have around 2,300,000 uh series uh concurrently and we have defined some on our parameters configuration some attention policies and we are using now PVCs so one parameters will

not lo the data again and then we have it also some um we are sending the data for some long-term storage outside of the cluster we are doing remote but for this we need to filter the data because we cannot send uh and then so it was supposed to be a demo time but I couldn't have my demo in time so it will be just slides but

what I will explain is the different dash the different dashboard that we have so the first dashboard will be an overview for the user cloud analytics and it provides uh visualization for the users about their workloads and how can they see and debug any issues with them. So as we can see here so this is the view for the user. Uh he can see we have the

average GPU consumption and memory and CPU and memory utilization per name space. So it's all scoped by namespace and people can look for their own workloads. And uh next to it we have a little bit um an overview of the available noise sources and the use sources for that name space. And we have some different uh graphs for the GPU metrics. And then if you scroll a

little bit down in this dashboard you have also the the other metrics. And here I will tell a very short anecdote uh about one of our users. Uh he was doing some uh quantum lats field simulation as one do obviously and um he was earning some benchmarks with that and the benchmark they were fine but there were some fluctuations on the scores at the end and he

was just using the same hardware and he was wondering what was going on like he went onto uh the dashboard and he was checking the different metics the GPU utilization was obviously 100% % all the time. The memory usage was fine. Also the CPU and we couldn't figure out what was happening. One of the panels that we have, I don't know if we have it here. No,

it's not here. But it was the frequency of the CPU and we figured out that the frequency of CPU was fluctuating a little bit and that was caused by the default uh kernel in Linux. he has the CPU governor policy and that policy was just modifying uh was modifying the frequency of the CPU. So when you are on the same workload you may not get the same

performance and we need to change this on the kernel level. So this is this kind of dashboards really enable people to uh dig deep in their problems and figure out what was happening. And then for the second overview, we'll have an a cluster and this dashboard will be uh a way to aggregate and all the metrics that we have in the whole cluster and we can also

the check the historical for capacity planning. So as you can see here uh we have some average averages of itization on the whole cluster with little bit of the number of users and the power metrics that we have and some GPU metics here and uh we can see also the different nodes that we have and one important aspect of this dashboard is the the panel at the

bottom which is the idle session and this is really important to have because some people will just allocate some sessions and they keep them idle and we needed to know this and then alert the users about this and it works pretty well when you are uh kind to people and send like uh professional emails they get the sessions and finally uh the final view that we have

is the sustainability. So we have a beautiful dashboard that was done by lawyer and it has all the power metrics and emission metics the CO2 equivalent and it's pretty good I would say and then yeah finally uh after all of this work there were some challenges uh along the way uh one of them is we figured out that sometimes when you put the new hardware in your

cluster and you have some old oss laying around, it may not work. And this was um some of the new AMD CPUs, they don't or the kernel doesn't support them for the rappel module. So, we had to figure out something for that. Another thing is sometimes some new features that are we were excited about can change some assumptions. So we are assuming that the tool that we

were using will behave this way and we are accelerated by a new feature. We deploy the new version and then it broke everything and D who is sitting there was coming to my office knocking at my door and saying just stop just stop. And the final thing is uh you need to pay attention for the MC partitions because sometimes they have their own particularities. They are not

for a full GPU. they are just in disguise. And uh for the outlooks we are looking uh to collaborate with the Kepler team to bring the power monitoring to the VMs because now it's only um back metal nodes but we want also to add this capability for the VMs and also we want to add some dashboards and visualization for cope flow uh service level metrics and then

I will let Laura present this >> Yes. uh as Amin was mentioning we're learning the hard way. We think that for you also uh configuring each of these component is kind of t tedious and errorprone and it can takes weeks of effort to bring build everything from scratch. So we wanted to give you a possibility to go from zero to all this observability in possibly less than

five minutes. And that's why we open an issue upstream. So if you could scan the QR code, put a thumbs up and everything, this would be really nice. So what we would like to do is to have a >> I I see a lot of phones up. Thank you very much. >> Please uh so to have a cubeflow observability manifest. So a simple recipe that will deploy

um a sane from a fuse config uh some uh all the needed service monitor as well as well like some pre-built config graphana dashboards and it will also have configuration about the Kepler and all the GPU metrics also that we are collecting from for example DCGM and so on and yes if you could work and make a thumbs up thumbs up it would be nice. Thank you.

>> Thank you Laura. And now some final words of wisdom. So the first one is don't let your users be in the dark. And this is is they will thank you for this. And we will thank them also because we learned a little bit of stuff by just giving them the capability to see all the metrics they want. The other thing is power aware AI is no

longer optional. We need to think about the sustainability and we cannot just build nuclear power plants with each data center even though some people think it's possible but let's think about the planet and then uh don't let me alone in the issue please get involved and the bigger sentiment also is get involved in open source contribute upstream open issues uh try new things we are all here

for the open source and what's next reality and thank you very much. Any questions? >> If you don't have any question, thank >> Just wait a little bit. Yes. >> Okay. I think no question. Everything was clear. Thank you very much everybody. >> Ah okay. We have we have Thank you very much. So it was a great talk. I'm just curious for you what has been the

most challenging in terms of observability because I saw for us we've been running QFlow for five years and I think that it's it's not obvious where to start. We move from graphana to data dog and then back and then we are like but from your point of view what is this like minimal viable graphana dashboard like I know we saw different things and I I also think

that the planet should be taken care of but if you could like recommend the minimal things what what would that be so where to start? Thanks. uh I I think the manifest that we will be providing so I opened the issue and we we are thinking about providing the manifest will be the place to start. So they will provide you with all the metrics or or like

all the configuration to have the metrics that you need and the panels that you need and it's it's a good start because most of the people that will be running uh workloads on coupubeflow will be ML related and the basic stack is not made for theirs and it's mostly focused on microservices. So what we we are proposing to contribute I think will be a good start >>

Hi. Um I wanted to understand your thinking behind tracking power. Um I think generally like the what people say is you want to maximize GPU utilization. You want faster training iterations and generally what that means is like putting more power draw on your GPUs and that's something researchers are rewarded for. So how do you sort of like balance you know wanting to optimize GPU utilization with wanting

to I guess like minimize your power draw? Uh I think the idea behind the the power uh giving the parametics for the users because physicists are trying different accelerator AMD basically and uh uh Nvidia GPUs and they are trying also on FPGAAS and they want to see um how efficient each hardware is for their calculations and because they need to make this efficient calculations they need to

put them later in the experiments and they are is constrained. So any gain in efficiency is good for them. But it's not if you consume more power, it's not it doesn't mean it's always better. It may give you a little bit of boost of benchmark, but if the power consumption is increasing more than the the gains in benchmarks, it's not worth it. >> I see. So you're

basically suggesting that like some GPUs types differ in terms of like the flops per like kilowatt hour that they can draw. >> Yep. And this is also based on the workloads they are running because uh like I said it's not only machine learning workloads but they will have some computing scientific simulation and stuff and they may behave differently uh other than the machine learning