K8s-sigs NFD × SYLVA: Declarative Image-to-Node Compatibi... Eduardo Arango Gutierrez & Chaoyi Huang
About this talk
This talk discusses the integration of Node Feature Discovery (NFD) with the Silva Foundation to address challenges related to heterogeneous hardware in cloud and Kubernetes environments. The speaker outlines how academic institutions and end-users experienced difficulties in deploying containerized applications across diverse architectures, revealing a gap in compatibility and performance. By enhancing NFD to produce metadata about the system into the container images, the team enables better scheduling of workloads on nodes that suit specific hardware configurations. This implementation involves creating an NFD client for improved artifact handling and scheduling, ensuring that containers are routed to the appropriate nodes based on their features. The speaker emphasizes the importance of managing deployments effectively in complex systems, highlighting collaboration among open source communities to refine solutions for real-world problems.
Full transcript
Hey everyone, uh thank you for being here and for everyone watching this uh on your homes in YouTube. Hi, welcome to this talk and let's bring us home, right? Like this is the last session at CUCOM. Let's make it fun. Let's make it quick, but also uh I'm super super excited to give this talk because this is a project that I really care about. I've been trying
to have the NFD awards in a slide at coupon for years. Uh many many thank you to Marcus from Intel who is not here with me and many thank you to all the other NFD maintainers that have taken this project over the years. Uh so this today is uh integration that we have done with NFD and the Silva Foundation which we will talk later on on solving
problems from sik cubernetes on real life. So this talk is about a real end user that uh they had a problem and we as a community got together and created solutions within a sik Kubernetes project to help them move forward and the problem is a very interesting problem and a very old school containers problem. We have the promise on the clouds that once you build a container
you can ship it everywhere and it will run everywhere. Well, there are some corner cases and not that corner like Telos are a big customer not corner case where they have tons of different SOC's right like they at your home every rout like Wi-Fi router has a different chip and if we go to Telos then every Wi-Fi router is different every antenna is different every 5G antenna
has different SOC's and chipping containers there is very complicated But the second use case that also there are a lot of contributors to this effort is university is academia and is HPC in HPC mostly because of the procurement process that happens within big institutions they end up having very clusters right so it's not a big cluster with a a singular architecture but they have uh chips from
10 years ago that are still running because magic exist all the way to new chips and shipping one container to that academic cluster can lead to some errors. So this whole uh talk today is about how can we address this problem of having very eterogeneous architectures within uh or helping with Kubernetes. So we are collaborating. So the cubernetes community got together to help the Linux foundation sila
which is another foundation within the Linux foundation and they had this big problem of having very heterogeneous hardware deployments. So representing the Kubernetes side it's me uh Eduardo I work at Nvidia mostly all things Kubernetes and you okay my name is Joe Ham and uh working for the Huawei and contributor to uh uh Kubernetes CNCF and SA and uh uh let me talk about uh Oh, you're
the next >> challenge first. >> Yeah. T as the society infrastructure, it has to provide uh 59 service availability. That means less than six minutes downtime is allowed per year. Uh in many countries, the government will set SLA for telecom services. for example cop plate and uh the network had to provide and a guarantee missing critic application requirement for example for remote lowot control less than 5
millisecond latency. >> Yeah. So network outage usually will lead to huge economic loss. Uh in the past there are lot of use cases like uh the network outage. Many people are impacted by the network outage and millions of people may cannot connect to the internet and the things are disconnected from the network. So unexpected uh network outage in some country will lead to penalty and the network
is consist of lot of interconnected network devices and the network outage is usually caused by some min issue. one minor issue will be propagated and enlarged to other network part. So we call this phenomena as the butterfly effect. Uh in the past the vendor is responsible for integrates network device from hardware to operating system to uh Kubernetes platform and the application and uh to meet this strict
te telecom services requirement for example uh 59 the vendor usually We do various type of test in horse. For example, uh more than 10 times traffic test to make sure the network device can survive from the traffic storm. And uh the vendor also may have to track every six core interfaces to make sure the telecom application is compatible to underlying hardware and operating system. So uh currently
operator expect the vendor to deliver their telecom application to third party uh kubernetes platform and underlying hardware and operating system. So the missing part uh here is that uh the compatibility between vendors application and the underlying operating system and hardware because the third party uh Kubernetes platform usually like IT system just provide the So the vendor need to deliver immediate uh with the telecom application together to
the operators so that operator can validate the telecom application is compatible to the third party kubernetes platform for example for hardware the third party hardware and operating system so container can one doesn't mean it work in telecom industry because it has to meet uh the SOA and the latency and other requirements uh to load out the telecom application into production. The low out wind window is quite
short. It's us you usually just four hours you have to done the upgrade or scaling in the night from 2 a.m. to 6:00 a.m. for example. So the low out the loout is done in so short time window there is no chance to de debug any issue. any issue happened have to lower back have to to lower back. So it's very important that to do the uh
compatibility validation before the low out and schedule report to compatible node. Make sure it must 100 guarantee that 100% guarantee that the node the application must run on compatible node. Uh SA is a Linux foundation opensource project. It was funded by five Taiwan Europer operators and two vendors. SA is to release Kubernetes software stack and uh for tailored for telecom application and edge use cases. So uh
there is also validation program in SA to make sure the telecom application can be validated against the SA stack to make make sure they can work properly. And silver is to provide the workload cluster automation and um with different currently there are two type of operating system. One is source and the other one is Ubuntu. And the workload cluster can be run on the infil structure like
bame or uh open stack or VMware based or open and there are several type flavor of workload cluster for example A2 and so on. So CNF is the telecom application. The validate program is to validate the telecom application. CF can work on the workload cluster. Uh, SA validation has uh nine validates platform in total and uh currently 22 uh vendors product validated in uh SA validation center.
So uh the image compatibility uh has already been integrated into stea validation center and uh it provide the automation script and tools to validate the uh application image can be compatible with the target node. It can use uh kubernetes job to do the validation. And there are also some community developer uh image compatibility artifacts. Uh but it's not necessarily to it it's not mandatory the vendor to
publish their image compatibility in public uh visible repo. We are respecting uh vendors confidential. Yeah. So next uh I will hand over the talk to Ado for the solution. >> Thank you. when the Silva Foundation reached to us it was uh by the hand of academic institutions and it was mostly because currently if you have used NFD it means that you have seen what NFD gets from
your system it advertised from your system and the the initial idea that was very simple. How can we get all the labels and information that NFD knows about the system into the container image? So into the OCI spec in such a way that when we move this image around we know where the image was built. So going back to the the Silva certification center for example. So
every time the image gets built for the infinite permutation of possibilities that you just saw. Then the information on where this image was built gets stored inside the container itself. Then when it being deployed, we build a specific uh artifacts so we can route the images to the proper node and this way get better performance and better uh SLOs's for for Silva. And what we have is
the the solution here is we build the image plus the artifact. And this artifact is basically everything that NFD can note from your node which is a lot for like it's a very big list from uh everything about your CPU your PCI and uh all or the node itself to the point that NFD can tell you if a USB was plugged or not into your server at
the time NFD was scraping information. Then we put this into the registry and we are using oras. So this is not just O the OCI registry standard but this is the oras standard which allows you to push not just the image but a full artifact and then in the cluster we built an NFD client. So before this collaboration, NFD was just a master worker architecture. But for
solving this uh problem with the Silva Foundation, we built an NFD CLI and then with this NFD client, we can look into the cluster what like what are the set of nodes that are better to run this container itself. And we also have admission controllers or gates that can help us route our containers inside the artifact. So for those that don't know uh oras is a way
u it's a kind of like a meta registry where you can store more than just plain OCI images but uh full artifacts right so what we I don't know if you can see it okay yeah I think it's visible so this is the jaml that will be packaged next to your container image and it will contain as much information as we can out of your node And
this information can later be used for a scheduling the image. So the rule mapping and node feature groups. While we were building the CLI, we also realized we needed a a second artifact from NFD. And this is the node feature groups. No feature groups allow you to create basically a catalog of your cluster or of your system. So you can create uh so no feature group is
a CRD and you can create multiple custom resources basically aggregating your infrastructure into a specific uh inventory cataloges uh that this is a manual step right like you can simply aggregate your infrastructure in this cushion resources and the rule mapping that then we have internally at Nvidia is as simple as we have here that is if there is a GPU in the system and there is a
region then aggregated in a in a specific group and internally NFD has a a full engine to read rules in reax mode in uh a standard mode in in yeah we have like four or five ways of of reading rules and then we can use these rules to map your node your pods into the specific group right so you can divide your cluster or your system Because
in telco I don't think you say cluster is is a system. It's a disagregated architecture into groups right like you can have the group of hardware that you can use for a IML the group of hardware that you have for high performance computing meaning like now that this uh AI is is on fashion it will be like the nodes that you have for training and group C
general purpose right so even though you have a big cluster a multicluster setup you can create a catalog or inventory of infrastructure using the NFD node Fisher groups some lessons and reality checks that took us here, right? The registry behavior was not as as we expected. So we had to navigate across the limitations of auras registries uh kernel driven and multi-burner silicon. As previously you saw there
is a almost infinite permutation of SOC's that can go into telco devices as well into academic HPC infrastructure where you can find a plethora of multiple solutions that have come over time for academic environments. So sometimes when you build a container for example a a clear example is in the academic side uh they thought that building a a build server farm for container was the right solution
and the thing is that this build farm was just with one specific chip and they will build the chips there and then chip it to the cluster that had many different permutations of set chip and thinking that just because it's a AMD64 for it should run and yes the application would run but we found like multiple differences in performance and sometimes uh even ABI incompatibilities and the
application will break and debugging that all the time led to a very specific hardware incompatibility even though it was for example the same uh AMD 64 from the same same year but it was a different architecture the overall So I'm going to later on demonstrate that you can create different specific groups using the the node feature group uh feature from NFD and how can you use that
tied to the image compatibility that we build into NFD to route your containers to the place that they need to go or as I can show here we can also have a polity failures where we reject a pot because it is trying to be allocated in a node that uh should not be allocated, right? So, Is this okay? Yes, no complaints. I think at in the end
I will have to shrink it down because I display a jaml and to show the full jaml I have to reduce the the so what I'm going to do here is to deploy a kind cluster a very simple kind cluster with uh two uh two nodes two nfd workers and I'm going to use the nfd uh node feature group to try to uh route a container to
be allocated where I want it to be allocated. This is already too small. Yeah. Here. So, I'm building the images and deploying kind Okay. So now I'm going to configure our kind cluster, deploy NFD and deploy the node feature group uh custom resource definition. So you can see here I'm deploying the node feature group API. So it is uh right now feature gated by default and is
not because it's on alpha or beta is because this feature uh is very specific for a specific use cases. So for NFD we will keep it uh off by default and you can turn it on as I'm showing here with uh with these flags when deploying it. Okay. Okay, so NFD is deployed is running and the node feature uh group is uh enabled. Now I'm deploying a
a small uh image compatibility. This part is new and will be cheap in the next feature uh next release of NFD. currently is being u is an open cap against NFD which we will have a a scheduleuler that uses the scheduleuler framework from Kubernetes to use the information from node feature groups and the image compatibility framework to route containers into the the proper So we set the
image and image uh pool policy all based in the image compatibility framework. Here uh I'm deploying the scheduleuler config. So for the cubernetes framework so it works with the customuler that we are going to have for NFD. And then I will try to push the images that I have locally here. So this requires this right. So what I'm trying to show here is that we can then
have a very fine grain control or where do we want to deploy our pots for when we have these very specific use cases that we want to control uh our systems. So we attach the OCI image compatibility artifact and we can use uh one tool is uh the recct artifact and this the artifact itself is generated by NFD. So this is just showing you that we push
the artifact to the local registry of kind. And here is an example of when you do when you deploy NFD and you do cubectl describe node NFD knows a lot about your system more than you will ever think and the things that inspire this talk for example is that you can see there are tons of different variables and sometimes for telos and SOC's there is a infinite
permutation of these being true or false, right? And even though you think it's just because it's AMD 64 or ARM 64, they should be the same. Uh these things could be different and lead to application crashes and failures. This is just a bare example of what NFD knows about your system. I'm just showing here the CPU information, the kernel information, and system information. Uh but again you
can uh tune NFD to show more information or even not to show information depending on what what's of your interest. For example, uh if you are using a GPU currently uh in Kubernetes, the GPU operator uses the information that comes from NFD to know where do you have GPUs in your Kubernetes clusters and then there do like create all the needed artifacts so you can use the
GPUs in the specific systems. So now I'm going to schedule the pod using the image compatibility. So I now have the uh the node labeled with NFD. I have the artifact attached in the in the registry. And now I'm going to use a scheduleuler that knows about NFD and knows about the artifact in the image. And there what the is a very basic operation is the scheduleuler
using the kernetesul framework knows about NFD and NFD labels and annotations and know that the images in the registry are attached with an artifact that has the NFD API right so since it's the same API so the labels and the information in the artifact share the same uh schema for the scheduleuler is really easy to match where this image should go and then as we can see
here it was running and it has a mash status right so NFD has a an internal masher engine yeah and we have here the no feature groups so we have a node feature group with this specific set of match features and cleaning up here. So as you could see it's it looks very simple but actually took us almost like a full year to get here uh because
all the different permutations the problems that we were facing uh creating the CLI and trying to respect NFDS schema uh was a lot of back and forth and it was multiple uh vendors and stakeholders involved in this and the good thing is that now we are going from what Joe was saying We're going from problems at 3:00 a.m. because uh you try to deploy the same image
to multiple architectures to having a way to reject if a image should not be deployed to a specific node. uh the benefits for the ecosystem are broader than just telos right like as as I'm saying uh we had some contributors to this idea from universities because this is a very uh important problem for academia and for universities where they have a plethora of uh different architectures uh
even from my university I know of sometimes universities end up using regular desktops as compute infrastructure so having a very heterogeneous compute ecosystem is really hard to manage. So the benefits to the ecosystem uh we can now have multi-bendor support in clusters and everything is done under a standardized cubernetes primitives meaning we build all of this solution on top of what NFD has been doing for years.
So it's a battle tested API and we are always looking for contributors to NFD. So please uh the scheduleuler that I just show is being discussed. It's an open PR. Uh we hope to have it for the next release. So it's still time for you if you want to contribute, participate and and review the PRs that we are having around thisuler. Uh it's still open and we
wrote a blog post on explaining this in in very specific detail. the blog post using the Kubernetes uh blog post uh web page and questions. >> Where to get my water? >> Thank you for the presentation. Uh the question is as far as I got your presentation there is a moment when people find that two nodes which was considered to be the same class are actually different
for a new feature which was shipped set of instructions new feature in hardware activated etc. And uh how does this postfactum discovery matches the requirements for having uh successful deployment. So uh if we have postfactum discovery this is like postmortm incident >> and it's it doesn't solve the problem from use case uh how you deal with this. >> Yeah. So the thing is usually as as Joe
showed at for example the Silva Foundation they have a certification environment right but sometimes we forget that in these certification environments the system admin enables or disables a specific kernel flags for example but these kernel flags are enabled in the production system but since this is so such a fine grain control flag sometimes we forget about that. So what we are doing here is that we are
basically keeping a full inventory of the system where the image was built and then we have the scheduleuler to look where is a node in your in your production system that mashes everything to the place where the image was built. And if there is no node similar to it, it will not schedule the pot. >> That's kind of happy path, a set path. You stopped deployment before
it broke. >> But often it's like a minor revision for hardware and uh before some hardware feature wasn't used and your application just start using some odd bit from hardware and turned out that revision 3 has this problem fixed. Revision two hadn't. And before you was considered the same ids the same nodes and suddenly they are not uh is there any mechanism to discover this diversification of
classes >> yeah so NFD discovers everything so that's not the problem the problem that you're saying of understanding backwards and forwards compatibility it can be man uh managed via the matchers right so what we have in the note feature group of NFD is that you can have matchers and you can have ranges of possibilities. So it's not like if it it doesn't have this kernel flag enabled
don't schedule you can have like room for error for send it like that right. So basically you use NFD masher to define your rules understanding your application like this this is very that's why like telos are a good example of this. This is a a use case that is kind of like edge and it's a corner case where like fine grain control can lead to errors. I
know this is not a problem for like the 95% of us all because applications are very flexible on this on this matter but for HPC applications meaning when you're running a big uh mathematical simulation that you know is going to take days one kernel flag can be the difference on the performance of set mathematical simulation being uh uh good or bad Right. So this is where you
really need fine grain control of things and for this fine grain control then we have the NFD uh matching system. So is when you really really want to go to the fine grain detail but for for the big of us conversation it's yeah I understand it's not a problem. I I hope I'm answering your question if if >> uh very interesting. Thank you. Um, so, uh, I've
never run into this, but I don't work for Telico. Um, but I have run into running a container like built for x86 on ARM and typic or there's a solution for that in the normal ecosystem to just make a multi-archch container. Is there support in this ecosystem just to you know and in that environment you build multiple containers and then you use metadata to like you know
essentially tag multiple images to separate containers. Do you support that with this? >> Every container that you build, so you're saying multiple containers, right? Every container that you build will be back with the artifact of where it was built. And then if you push multiple images to a registry, each container will have their own metadata of where they were built. And then using the node feature group,
you can route them. So yes, we support like multiple uh container architecture for building the same application on on different uh and that's the whole idea right like that is kind of like the the whole idea of this is that you will build the same image in multiple architectures host them on a registry and then using the node feature group and using the NFDSuler uh plug-in you
will only pull and route the image that you want that you care >> I guess my question is that metadata building with the image in that registry is the same registry that you would use for a normal multiarchch build like if you use build x. >> Yeah, so supports uh multiarchch images. >> But is that the same registry that you would get in like DoctorHub or is
it a different registry? Does that make sense? >> Yeah, it's a Oras registry. Okay. >> So, but I think DockerHub now supports ORAS API. >> Yes. >> But yeah, if no more questions, uh we can wrap up CubeCom. Thank you everyone.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32