Solving Industrial Challenges With KubeEdge: A Post-Graduation... Yue Bao, Hongbing Zhang & Yin Ding
About this talk
This talk focuses on the KubeEdge project, which aims to solve cloud-edge coordination problems using Kubernetes control planes. The speakers, Inding and Honging, highlight their journey from starting the project in 2018 to its graduation in October 2024 as the first edge project in the Cloud Native Computing Foundation. They discuss KubeEdge's capability to handle various industry challenges related to latency, unreliable connections, and data privacy, particularly through federated learning and joint inference methods. The session covers several use cases, including applications in transportation, healthcare, and smart retail, showcasing KubeEdge's ability to enable efficient edge computing through AI and cloud collaboration. Additionally, the speakers emphasize the importance of community governance and partnerships in growing the project's ecosystem.
Full transcript
Hi, welcome to our maintenance track. Our today's topic will focus on solving the industry challenge with Kubage. So, uh we our project graduated on 2024. So, this one we're focusing on the recent AI focusing and also a lot of a case study. So, uh this is my colleague uh Honging. He is a TSA member. My name is Inding. I TSC member of uh Kuba edge. So we
started uh this this project about 10 years ago. So it's a long journey. I would like to share our journey with you guys today. So first uh you can see in uh keynote yesterday uh we our project are very honored. So it's mentioned this good slider is behind the kubit deployment. So it's based on kuba deployment. This a our sixth time we are on the kubcon keynote.
So very excited notes. Uh thank you everyone. So we are really excited to power industry user cases behind by our kovage project. So uh we are a innovated project journey. I will skimmed it. So it's back to 2018. We are part of a uh Kubernetes IoT work group. We launched it on the November 2018. Then we donate this project to CNL Foundation on 2019 uh March. And
it's a long journey. We uh become a became a uh incubation project in 2020. Then uh we finally graduated on 2024 October is on Chicago is Chicago Detroit uh is Chicago Kubakong. It's finally we graduate as a the first graduation project on the cloud native edge. So no other in the CNF blueprint there's no other edge project graded. So we were the first one. So I will
quickly introduce our K edge basically is we want to use the Kubernetes control plane to help solve the cloud edge coordination pro uh the problems. So usually this case is different because u it's cloud edge the deployment scenario is different from pure data center deployment. So we solved the problem from the cloud and edge there's a long latency a date connection is not reliable data privacy all
these things we try to solve by uh cloud edge code coordination problem by kub edge basically the basic idea was just deploy the control plane on the cloud and not the traditional kuballet not the cloud the worker node not in the cloud but in the edge side. So from cloud edge we set up a new duplex the two communication channel to secure cloud communication and the uh
synchronization we are solving the uh long latency issues and the unreliable connection. So if the connection interrupted, we can restore edge node back to the desired state controlled by the control plane. So uh currently we have more than 8,000 stars in on the GitHub. It's more than 2,000 uh forks on the GitHub and totally is about uh 1,600 contributors from 100 plus organizations. So it's thank you
for everyone contributing to our So I quickly introduced the our sig AI effort. So uh in kubage we have a few six we have a s uh networking signal uh sik aai sig rob robotics. So this one I want to just quickly mention because it's only 30 minutes session. I just quickly mentioned the CGI that's was we uh support all the main stream of AI frameworks including
tensor, pyarch, pedal pedal from bu and man scope man and so we support uh joint inference increment learning federated learning I will go a little bit detail on this one and also um trying to see the data and model management across not only on the cloud but cloud and edge. So for uh thing is we call it feder federated learning is different from uh the traditional traditional
u combined learning. So basically we want to isolate the data for data privacy. So the sensitive data never leave your edge. So that for the enterprise use very important. So we can have a edge side learning and we aggregate all the intermediate result back to the cloud to consolidate uh based on the inter each individual training result. So we call it federated because we combine the result
together uh because each edge node may have a big uh variations. Some are really tiny and some are really powerful and also your data set. So some edge may have only a few southern date point. However, other huge edge may have a gigabyte or even terabytes data because um we had some examples from the hospital some only a small for example in the hospital uh scenarios. So
some they deploy on the simple clinic only have a few patient data. However, in the big hospital you have a I mean millions of us patient data. However, all this data are have a privacy issues and very sensitive. So it should never leave the local facility. So we try to train this data locally and on the cloud we aggregate the data because each model will train with
different number of data set. They have a bias. So we try to uh solve that together. If you are interested in the detail, you can come to our ad discussion. We have a regular meeting every week uh to discuss the this federated learning issues. The other one is called joint inference. So this focus on uh the previous one is focusing on the training. This one is focusing
on the inference. So basically for you know usually the edge is have a constrained resource so they cannot handle the big model. So we try to have we call a joint inference. So we deploy uh compress small model on the edge side and you have your comprehensive big model deploy on the cloud. So on the edge side we are going to use the small model to do
the inference. If result reach the threshold so it's a satisfy result we just return to the user directly. However if this result not good enough we just upload this data to the cloud to use a big model to do the inference to run this on the large model. So have a better data. So basically on the edge node you have the small model do the shallow shadow
model to the quick inference and in the on the cloud you have a deep model to do the comprehensive inference. So uh now I pass to my colleague uh Ping to do the case study. >> Okay. Thanks. Uh yeah like you introduced our cuber edge got graduated you know two years ago. I think uh graduation for a project is not easy. Yeah, under CSF umbrella only about
20 to 30 project got graduated. So I think for cuber edge it's a graduation it's not just because of its tech technical maturity but also a lot of end user and case studies. You can say cuber edge as edge computing it can support various industries and various scenarios. So here I can list a few industries and the scenarios we supported in past 10 years. So you can
see we had a big coverage right for the industry we have transportation, energy, industrial, manufacturing even CD and vehicle satellite right and the financial logistic a lot of uh industries and customers. So as you can see from this morning's uh keynote we have a uh no earth's orbit is a satellite actually satinite was using kuber edge to process data several years So here I can give some
details about our you know uh very influential customer case how we can enable cuber edge into satinet for edge computing. So the background is so right now there are many many senets in the sky right but sadet has limited the computation power and also the network bandwidth between satinet and the ground station is pretty limited. So as a result you can see every day the senet will
capture connect and produce much data but the majority of the data cannot be transferred back to ground station for further processing. So that's a dilemma right how to balance the huge amount of data and the limited you know computation and the network bandwidth. So the solution is we can make sadet as a edge node. The application or the AI inference will be done in the satinet. So
you can see from the picture right. So the sadet is running as a cuber edge edge node and the ground station can be as a cloud So the cloud and the senet should be cloud and the edge collaboration. Right? So when that data was connected and captured on the sadet the sadet will run an application or an AI model to process the data. So that means the
data can should not be transferred back to only processed result can be transferred back. So as a result you can see the data was processed on the sedanet itself only a small amount of data or result can be transferred back to ground station for monitoring for further processing. So this is a very typical scenario and this was a very influential customer case and it was u you
know mentioned by many kynote speakers in several kubric comps in past you know 10 years. So this is a very typical case for the cuber edge is a cloud edge So the second one is also a influential customer case is cuber edge to enable and power expressway system especially for the toll system to gate. Uh this is a real case in China. You can say in China
there are many many maybe more than I think um 100,000 heterogeneous node or targetate in China right it's a big network the reality for the express gateway target is the first every to it has different devices some device are s x x86 device some device are you know based architecture even maybe a further risk file. So the device is heterogeneous. How to manage heterogeneous devices is one
problem. The second one is the amount of the devices were huge. It's about 100,000. Right? How to operate and maintain easily or efficiently should be one problem. For example, how to upgrade the application? How to upgrade the firmware? If for the traditional method you you know burn and download and burn one by one that mean cost a lot of time right and you know resource. The third
one is for the target actually is edge device or edge computing device besides for the you know toning system it can run other applications for example there there there has one video surveillance in every target right how to manage and how to control this video streaming is not just for the you know the tall system so the to solve this challenges we can use cuber edge. Firstly,
we can manage 100,000 heterogeneous with different architectures and we manage about you know 50,000 applications and we can have the easy operation and the maintenance model for example we can one button to upgrade the firmware right to upgrade the whole system uh we can have the upgradation of the new applications so as a result this will be a big network of the the Hannah Express highway to
system. So this is another typical scenarios and use case we used cuber edge in this you know big and system at scale. So the third one is also for the new energy vehicle uh we can say cloud native based vehicle also this customer case was introduced as a keynote speech in uh last year's uh uh kubric China so in China we have many many you know new
energy vehicles right and we have many upstream or mainstream applications which are developed by using cloud cloud native methodologies for example the CI CD or microservices even managed by you know cloud-based kubernetes clusters so every car should be a edge device uh which is running various um applications or AI models uh so this will have a big network right so when the cloud system had developed one
application it will have some you know maybe s send it to some cars which are you know maybe exper firstly uh experimental cars then with some testing we can you know push to many many um you know real cars right so on the right side we listed the one typical scenarios uh which is to predict the life cycle of the new energy cars battery you can see
the battery as uh in the car as edge device the firstly the life cycle the lifetime of the battery should be inferenced by AI model. So when car is running many data running data was captured by the car. So the car side of the edge side will do some feature extraction from the edge side then send the process the data back to the cloud side. The cloud
side will launch an AI port to inference if the model is not accurate enough. So the data will the the cloud side will continuously training iterating the AI model. So this is a closed loop right. So the cloud cloud side for AI inference and the new versions AI model training for the edge side is to capture the data for you know simple feature extraction or you know
data processing and for you know uh for cloud to further process that data. So this is also uh one typical case um for cuber edge to power a new energy vehicle. So the next one is for the smart retail it's also scenario based uh you know for the retail there are you know more than one one many uh retail stores right so every store will be controlled
by one edge system maybe one edge server uh which connect to many devices maybe some MQTT devices maybe some you know cameras right there are many scenarios for the smart retail this is one real customer case or real scen scenario. For example, when you are maybe you want to buy something, when you go to a store, uh when you stand in front of one, you know, a
shelf, maybe you are, you know, looking at some goods. So, there is a a vid. Oh, I'm sorry. Okay. Um when you stand for a few second in one in some in front of some, you know, shelf. So there's one camera who can detect your you know duration maybe they will detect I maybe you you have some interest for some goods right so this you know uh
this process was processed on the edge side when the cloud detect maybe you have some interest in some goods so the cloud will broadcast one stream of video so you can see the adverment or detailed description of this one of this good. So this is also a cross loop. Actually this business was you know widely used in some retail stores to to be more smart smart and
intelligent. Okay. Last but not the least also Cooper Edge embraced the embodied you know AI or the the uh the robots. Actually this video was also showed in the keynote uh every day because it's a typical scenarios uh for cuber edge to manage the you know collaborative robots. This scenario is pretty simple. Uh if you want to operate or control a group of robots, maybe you are
using with your you know natural language, right? Um so maybe when you're talking or go there um but you know the robots cannot understand the natural language the NOP with natur language model should be processed on the cloud side right for the cloud side understand your command and your attention of your of you. So the edge devices will translate your willingness to the command and it can
control a group of robots right. So you can see this your very small video. So you are talking in the chat box. All right. So give some um command and the uh edge side will calling the API of the cloud to for the natural language processing. they understand uh how to do you know with a result the edge device will control a group of robots uh this
group of robots are collaborative right and they will plan the plan the route and for each robot for that >> yeah okay Yeah. So next I will spend a few minutes to introduce how we achieve um this thriving community. So you know when I think many of you are running some open source project you know but if you want to get one project to be graduated it's
not easy right? So how to you know build a thriving community I think is one experience we can share to many of you. So first uh if you want your project to be graduated so first your project governance is pretty important. I think that's is a key point for TUC uh to evaluate and decide whether your project can be granted or not. Uh so you should have
a very open governance structure. Uh here we list our cuber edge community. We have development governance technical uh ecosystem and the and the user ecosystem. Even for the developed community we have uh two TSC uh which is the highest decision uh for the project and we have many things right for every AR sik robot sake security s AI and the device seek so right now we have
distributed and open the governance structure that means this project will not be controlled or dominated by one or few country or few companies it's really you know run by the community. So the second one is we had very good partnership across industry, academia and the research. So here you can see that many universities and many you know uh companies or many industrial uh researchers uh join our
cuber ed you know uh community. So they can provide a various uh contributions or you know scenarios um u beside uh the technical contribution. So third one is we launched a few engineering verification or interoperability either on the software side and on the hardware side. So this is you know similar with uh uh uh CNF uh AI conformance or you know similar program. So we can uh
have some verification for some software vendor integrators or hardware vendors uh to pass our verification. So I think this is pretty important if especially when we get graduated. So we will shift our you know uh running model from developerdriven community to end the user or industriven community. So the industry and the customer outreach and the promotion is pretty pretty important. So that's why every time we will
try to you know um uh uh get the opportunity to promote cuber edge in every kubriccom or even on the kyno speech right so we launched quite a few meetups uh end user uh scenarios or uh uh meetings and we also cooperate with universities to on the uh even on the university u uh promotion and we have uh you know grant a few awards to the developers
and the users to build a very you know extensive community also the uh mentorship program I think the CSF emphasize the mentorship program a lot right we uh after our graduation we every time we will launch quite a few mentorship program so if you are interested so welcome to join us right and we have you know continuous communication and the development. Um so the last but not
the least we have the very friendly uh developer portal with tools and we continuous to grow our uh developer community. Okay. So that's our you know uh sharing. So we here uh we listed our you know uh GitHub and the snack link and our goal and the vision is we want to make cloud native ubiquators. Okay, thanks very much.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32