KubeVirt's Evolution: Governance, Features, and Community Growth - Sreeja Varnam & Luboslav Pivarc
About this talk
This talk introduces KubeVirt, a project that integrates virtual machines into the Kubernetes ecosystem. The speaker explains the evolution of KubeVirt, highlighting its transition from a chaotic development phase to a more structured and mature community approach. They discuss the importance of establishing processes, such as the proposal and review mechanisms for features, which enhances transparency and predictability. A specific focus is placed on VM pools, a feature that allows for scalable management of virtual machine instances within Kubernetes, effectively filling the gaps left by existing automation methods. The speaker also outlines the key design aspects of VM pools, including their management strategies and auto-healing capabilities. The session emphasizes how adopting a structured web process fosters community trust and collaboration, ultimately leading to stable feature rollouts.
Full transcript
Welcome everyone. Uh so my name is Lo Slapage. I'm a software engineer and a dreadhead and cuber maintainer. Uh who was supposed to be here is it holder. Unfortunately he couldn't make it so he asked me to present it for him. Hopefully I will be a good enough replacement for you and let's start and of course uh I'm joined by here by Sria here. >> Yeah. Hi
I'm Sria uh working as software engineer at Nvidia. Uh yeah, this is my first CubeCon. Feel honored to be here. >> Yeah, thank you for joining us. Uh so let's start with an introduction we are uh Kubert is a project uh bringing a virtual machines into into the ecosystem of Kubernetes. You probably know that if not uh it's uh basically running on the Kubernetes. It has a
pretty much the same architecture as a kubernetes but we are doing virtual machines. It was uh it is about 80 years old. So we are in the ecosystem for some quite time and we are graduating this year u hopefully at least that's our plan right you can scan the QR code if you want to more know like what's the stage I think we are in the stage
of dual diligence so hopefully that will not take too much we had the first cubin summit in person we had many uh cub summits but only online. So I finally got to meet a lot of people in person. I actually saw their faces not only the screens and the voice and didn't only hear their voices. So if you haven't been there uh please watch the recording. There
were five talks very impressive talks very advanced talks. So if you are looking for inspiration kind of there's the good source of inspiration there. So every project gets uh uh through some kind of evolution and cubid is not uh any other project right. Uh every project starts uh from basically being as a child. Uh the child is very very much a self-centric. It's uh firing a features
uh as fast as it can. It's uh like the direction is probably um coming from one person or maybe like a group of small group of uh maintainers and it uh it's you know very chaotic and there is no really plan there is no really organization. So as you go you like kind of get more contributors on board because they think that oh this is a cool
project right like virtual machines on the Kubernetes or maybe they will like tell you like oh this is crazy but I like it as well. So they will join you and you are like kind of toddlers playing around still uh firing features every every like month for example and everything is good everything is happy but uh unfortunately as you scale more and more contributors are coming and
you got you got to start uh getting like some kind of users you need to uh mature a bit because the toddlers uh they kind of don't know how to behave to each other so sometimes they are not good to each other and maybe or not maybe I mean they don't really have a planning skills uh at that at that stage and when you get a or
when you want to get some good community and user base uh they all like the all the time the question is hey what's the road map like how is this looking is is the security working is it stable does it scale right and you get to the point where you cannot keep up uh with it if you don't have any process. we kind of matured um I
want to say like three years ago where we did like a official GA 1.0 release, right? And um that was cool. We got a little bit of uh slowdown like we didn't do a release every month. we went for Kubernetes schedule basically following it uh we got some process in place but it wasn't too much and uh the contribution contributors were not as wide I would say
and there was a couple of companies but um yeah some traction not too many not too much uh we got into uh present basically where we are kind of like I would call it mature community where a lot of other human beings which are adults uh are joining us and uh for this to work you and I you need you need a process right so let's let's
have a look at the process like what what we did for this to work together so uh we had to come up is a road map that's also requirement for graduation and uh we had to also get some uh more pro maintainers u so it will not be only few folks uh working on it and we for the process we choose to follow again kubernetes if you
didn't know we have like a concept in cubert is a razor it means if something is there don't reinvent it like bring it on and because we follow kubernetes so closely We wanted to follow the process as well. Uh and of course we didn't want to be too familiar to to get like confused when you would say like cape uh because it it could be cubert enhancement
proposals. So we called it virtualization enhancements proposals. So we proclaim ourselves as a only virtualization solution maybe a bit we are now uh so what what is the process about uh it's there to to give us a focus because we are maintainers and contributors from from so many companies but also individual contributors which are not attached to any company and we kind of need to know where
we are heading right uh so the uh The designs are there to collect like user feedback, user requirements, you know, what companies want to see as well. And once we uh once we get the submit of of these webs which looks like basically opening an issue describing a use case and then uh filing the the design as a as a PR into the repository. Then you get
to the review process of these of these webs kind of getting the feedback from our side. Is this something we are looking for? Maybe there is a better solution or you know you need to just improve a bit on it. Once we get there and every everything is fine, we accept the web. uh we collect few of them and we need to prioritize them right and when
we when we do the prioritization we are looking at the like kind of like what's the effect what's the impact and who wants it as well and how nice they are of and we don't want to anymore and and what I wanted to also say is that for the issue Why do we want to actually have a issue in inside the repository is for tracking purposes. It's
a it's a two side uh sides of the story there. One is uh we want the users to watch out for these new features and how they graduate inside the uh cubid but at the same time for maintainers and for approvers they need to track these features as well. So it will have a bit of technicalities inside it and also a bit of uh user feedback on
the issues hopefully and uh you should be able to track every basically change and our major or major discussion on these issues. So we have the issues, we have the designs. Uh let's talk about the ownership of these designs because if you have a community but nobody really owns anything, is it really something like does it work? Probably not. Um so the ownership is pro basically the
same as in Kubernetes. one or two or even more uh six or special interest groups are owning this uh process and and it can spine uh it can be a a bit of challenge to find the right sick but uh most of the time we are there to help that's pretty much it but uh last thing I want to mention is that um this collection of all
the webs is basically serving as a as a road map. So um how does this look in uh in perspective of our release? So as I mentioned already we are following the Kubernetes release which is three times a year and uh yesterday we got a release out it was 1.8 uh is there anyone who already tried out? Okay, maybe maybe you should. And so we are at
the stage of the beginning of the new cycle where we take about five weeks uh to look at the new features, new webss. We get them the most attentions uh from all the PRs on in the repositories and uh we settle uh with enhancements freeze which basically means okay we are done and for the next release we cannot take anymore. Um after that we uh transition to
the like develop development or implementation of these are driving the most of the most of the focus and most of the work and reviewers are focusing on these implementations. Um while we are doing that we are releasing an alpha bet and beta attack and release which can be then consumed and uh have a look at it and give a give a feedback for the feature which was
implemented. Unfortunately most of the time the features are getting very late in the cycle. So you are uh you don't have such a like chance to actually try them out. And we finish this phase with the code freeze. Uh when we reach the code freeze, we have a look at uh like okay is the is there was everything done. Do we need uh one more week for
maybe the work to finish and we get into the last phase with stabilization where everybody gets uh the time to do a bug fixes um you know improve CI improve flags and and that sort of thing. In this phase we do a release candidates at least two and if everything goes well we promote the release into release candidate into release. So there is a just just a
note about cuber calendar where everything is kind of settled. Uh there is also a release milestones where they happen uh also where we have a community meeting where each sik has a meeting. What I want to uh basically highlight is we have also a kind of web open door meeting where you can bring your ideas or a web and we try to help you to progress it.
And so let's have a have a look if this actually worked. Um before that uh what we wanted to achieve and what we think uh we achieved is three key points and that is transpar transparency, predictability and planning. Okay. Um some notable changes and uh are uh for example the DR integration from Nvidia or uh from Microsoft multihypervisor framework which uh we are slowly getting in and
a slight bonus is also that we are trying to use AI in to guide us through this process and maybe uh external contributors. Okay. And now the the stage is yours. >> Thank you. So uh basically I'll walk through uh a VM pool web and how web process helped us move the feature virtual machine pool from uh alpha to beta. So before diving into the core problem
uh that we are trying to solve uh let's go through uh the core concepts in cubert uh VM and VMI. uh understanding how VM defines the uh desired state and VMI represents the running instance. Uh these are the key to understanding how VM pools are built on top of them. So once we understand these fundamentals uh the problem statement and the design decisions uh behind the VM
pools uh become much more easier to follow and then we can go through the VM poolool what VM poolool actually is uh how they work and uh how the structured web discussions from proposal to uh reviewers feedback made the beta rollout uh far uh predictable and stable. So in Cuba there are two core concepts. Uh one is VM. Uh it's a Kubernetes resource uh that defines the
desired state. Uh how the VM should look like and how it should behave. Uh basically it's a higher level uh stateful definition that controls the life cycle of the VMI and there is VMI which represents the actual uh running instance uh like pod for VMs. uh each VMI maps uh directly to a uh live ku or KVM process on a node. So even if a VMI is
deleted or crashed uh VM definitions is still there. So in a simpler terms you can say that VM is the definition and VMI is the running uh live running workload. So the execution path looks like this. Uh so when a VM is uh created, cubert converts the definition into a VMI uh which represents the running VM and then word launcher part starts the VM using KU as
a user process and this allows uh VMs to run inside Kubernetes just like pods uh but by uh but backed by the KU or uh KVM virtualization. So what if you want to create uh multiple similar VMs? So the first obvious solution would be to uh create wrapper scripts or build automation to handle this 100 API calls. But uh the downsides with this approach is uh manual
scripting introduces uh inconsistency and it increases maintenance overhead. it also becomes complex uh errorprone it's hard to scale and maintain and when scripts drift and break when scripts drift uh it breaks with the API changes and there is one more solution uh which is virtual machine instance replica set uh while it uh creates and keep VMIS around uh but these only support stateless VMIS uh and cannot
manage VM It has a sophisticated uh limited sophisticated rollout strategies uh that are not suitable for long lived uh VM fleets. So this is where VM pools come in uh filling the gap uh left by uh automation scripts and virtual machine instance replica set uh by offering a native Kubernetes uh way to manage scale manage scale and maintain a group of uh similar VMs. So talking about
native Kubernetes uh uh if you look at patterns in Kubernetes uh a job goes as answering the declarative statement like uh uh create 100 objects. So it does that once and it's done uh it does not keep those 100 objects around and there is this replica set uh which is built on that. So it says create 100 objects and keep those 100 objects around and if any
of them disappears uh it recreates them and there is a stateful set that goes uh one step further like it says create 100 objects keep those 100 objects around and manage them carefully so VM pools were designed understanding these patterns basically it's a mix of a job replica set and a stateful set I mean it has a batch style API of a job, management of a replica
set and a stateful control of a stateful set. What you can do with VM poolool is uh you can scale out of VM instances based on utilization. You can scale in VM instances uh to optimize the cluster resource optimization. uh you can roll out uh changes uh in batches across a pool of uh replicas and it has the ability to automatically detect the misbehaving VMs and uh
spin up fresh new instances as a replacement and if you want to debug any uh one of the VMs uh you can normally detach the VMs for debugging and the missing VM will be spin up uh and replaced by the pool without uh automatically. So what you cannot do with VM pools is you cannot manage uh VMs which are dissimilar to each other. Uh basically it manages
only VMs which are similar in shape uh to one another and are derived from a single config. So let's dive into the VM poolool API and uh explore some of the key features uh that make VMOL a powerful way to manage groups of uh virtual machines. Uh first there is replicas uh which defines the number of uh VMs that should be running in a pool at any
given time and the controller ensures that uh uh the desired number of VMs are always maintained and there is max unavailable uh which represents a number of uh VMs in a pool that can be unavailable uh at a time during an update. Uh there is virtual machine template uh which describes the configuration for all VMs in a pool and every VM created by the VM pool follows
this temp template and there is another important feature that is the update strategy. Uh it defines how changes to the virtual machine template are applied to the existing VMs in the VM pool. So it doesn't just control how uh updates should happen like if VMIs are restarted or updated in place but also provides you the full control and predictability over how those update updates uh roll So
uh with update strategy uh we can decide which VMs are updated and in what order. So this is done through selection policies and ordered policies which are a prioritized list of rules that determine the update order and there is uh similarly there is scaling strategy defines how uh VMs are selected when the pool size is reduced. uh it also the it also follows the same selection policies
and uh ordered policies uh to ensure controlled uh and predictable scaling behavior and uh there is another important feature which is state preservation uh during scaling it can happen in two ways one is offline so when a VM is selected for scaling uh its PVCs are preserved instead of uh being deleted so that uh these PVCs can be re uh reused uh during scale scale out reducing
provisioning time and there is the online state preservation uh which is not implemented yet uh but the idea is to preserve not only the PVCs but also get the uh VM's memory state snapshot so that this allows the VMs to resume faster uh reducing both provisioning time and also the boot time. Another important feature that makes VM poolool uh reliable is uh autohealing capability. So it works
with startup failure threshold. Uh it defines how many uh consecutive startup failures are tolerated before a VM is uh considered unhealthy. So once it crosses that limit, the pool recreates the VM. So let's take a a real life example. uh let's say imagine a company with uh 50 employees and all these employees need uh identical virtual desktop environments to work. So instead of creating 50 individual uh
VMs manually now the platform or IT team can create a single uh config in keyword and use a VM pool uh with replicas as 50. So let's say on a fine morning uh one employees VM crashes in this case no manual action is needed the pool automatically uh recreates it because it has to maintain the uh desired state as 50 and if the new uh and if
the new VM that is recreated is also having few issues uh the platform team can come in and detach the VM uh from the VM pool and use it for debugging and and the new VM is recreated by the VM poolool that can be used by employee. And if you want to uh patch a security update uh basically to apply the base image based on the uh
based on the update strategy defined uh the VM gets updated. So this level of control allows platform team to perform safe and controlled uh rollouts avoiding unnecessary uh disruption. Uh let's say after a month or so a few of the employees uh left and the team size now becomes 40. You just change the replicas to 40 and the scaling strategy defined uh decides uh which VMs are
to be removed. So this makes the life of a platform team easier. Like instead of managing dozens of virtual machines, the team manages one declarative uh resource and the employees just log in and work. So a virtual machine pool uh is a strong example of how a web process help translate real user needs and uh needs into a productive ready feature. So the web process uh introduced
a much needed structure. Uh it clearly defined the problem we were trying to solve and uh enforce the early design reviews uh and help the community converge on the right API and the controller behavior and uh clear boundaries from the start. So by clarifying the problem early and aligning on the design up front uh the web process allowed us to move uh smoothly from design to beta
without major surprises and uh helped in getting timely reviews and feedback. Uh so in the end uh it's uh this isn't just about building a feature but uh it's about building in a way the community can trust and operate confidently and virtual machine pool shows uh how the web process uh exactly enable that. Uh with this we conclude. Thank you. Are there any questions? >> Question. >>
Yeah, go for it. >> I'm not an expert in cubit, but how do you monitor the health of a virtual machine? >> Uh so we as a pot pod has a readiness props and livveness props. uh the same is implemented for the v virtual machines and >> is there an agent that runs within the pod or >> you have multiple uh props you can have a HTTP
probe you can have a gas agent the gas agent is running in the virtual machine and we we are pinging it um yeah and pretty much that's it I think we have one more but I cannot remember >> in VMware you have the VMware tools so so this is something like that >> yes yes exactly the case agent is from cuma project and it's basically gent agent
communicating you you do a lot of stuff with it. For example, you can also uh freeze the file system when you do like a backup volume snapshot. Uh you could also execute some kind of commands and many more features. >> Okay, thank you. >> Hi, thank you for this presentation. Um I'm very interested in the in the journey that you had in term of project because we
we mostly have some devops engineers developers not interested in managing projects. So I'm interested in knowing for you how difficult it has been for developers to do this administrative job of maintaining a project. Do you dedicate people to that? Do you have ceremonies specific ceremonies in each sig and so >> yeah uh so I have a hope that it's not the overhead because if you have if
you want to have a striving project which will last you want the design so the design in my opinion is not the not something which is overhead but something which force you to think about the future the stability about the fe about the feedback about the motivation of course the writing is the overhead Um I have a hope again that uh if you write enough you get
better as with anything else in the life. Um and you and so Sria what's your feedback on on on this was was it the overhead uh if we haven't had this pro uh process do you think it would be easier? Uh no actually this process made uh our life easier like you go to community meeting you uh I mean provide what you wanted to do and you
get early feedback so that this made uh uh what we wanted to do made sense to all of them and they supported and made give uh helped us getting reviews and feedback early so that you don't have to uh wait for the reviews while after implementation. >> Yeah. And and for for the ceremonies, I mean it's open source project, so everybody can contribute and mostly the maintenance
are trying to keep up the project uh somehow in in a good way. Um we are just uh making up the process as we think it's it's the best suit for for us. For example, we just introduced the web open door meeting where you come and and talk about your web and if we if we see a need for any other like meeting or ceremony then we
we just add as we go basically. >> Does it answer your question? >> Yeah, totally. Um but still the load of managing of dealing with this project management topic is uh on the shoulder of maintainers who already have a lot to do in maintaining the project itself. >> Yes, that's true. But uh you can always ask for help in open source from others. So we are trying
to uh push it down to this uh six this special interest group as much as we can. and make the process like independent from us. But it's it's a slow process >> and of course we have for example 72 features coming in in the next release. So it's a lot of a lot. Yeah. >> It means also that SIG members are not maintainers. >> Uh well we
have a like kind of weird definition where we don't call them maintainers. We have like we call maintainers who take care of the whole project. But I would say if you are approver if you are uh then you are maintainer as well because you maintain the code right maintain the subsystem. So yeah it's a bit of terminology. Thank you very much.
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32