KubeCon + CloudNativeCon Europe

Longhorn: Intro, Deep Dive and Q&A - David Ko & Divya Mohan, SUSE

30:44 · 23 Mar 2026 – 26 Mar 2026 · YouTube

About this talk

This talk covers the latest updates and features of Longhorn, a cloud-native distributed block storage project, presented by David Co and Diva. The speaker discusses the project's governance, recent release cadence, and technical functionalities, including data engines and volume management. They highlight the new version 2 (V2) being developed for improved performance over version 1 (V1) and outline the objectives for achieving graduation within the CNCF ecosystem. Key enhancements such as better replica rebuilding and improved performance benchmarks are presented, along with community engagement initiatives to encourage contributions from both existing and new members.

Full transcript

Hey, hello everyone. Uh I'm David Co, uh engineering director at Susa. Usually I will do the one year's update every time in QCON. So um Diva. >> Yeah. Uh I'm Da and um I have recently taken taken over the community management for Longhorn um in an exclusive capacity. Uh so I'll be talking about uh what we're doing in the CNCF ecosystem to advance um our levels and

a bit about the governance and everything but um I'm sure all of yall here are more interested in the technical know-hows and you are ready to bombard David with the technical questions. So uh David please take it ahead. >> Yes. And we only have uh 30 minutes >> 30 minutes. Yes. So um this time let's general agendas uh because people may be here still not familiar with

long so still have to go through some overview and project and release updates especially we come out with a new release cadency I want everybody knows if you are planning to adapt longhome and data engines we have a different data engines right now and people ask ask about when is the v2 going to J and I will update some status here and performance matter So again and

lastly talk about the 1.11 key features we just deliver and upcoming 1.12 road map and this time is a little different we have a diva here to have some community update because we are planning in the shorten we want to move the graduation this our goal okay uh long overviews uh long obviously is a cloud native uh distribute bar storage hypercon converge storage together with your uh

wen nodes or even you want to do a comp uh control print node togethers for the testing or develop purpose. It provide a highly available per ving. Uh this is a fundamental part the interface for your workloads but we care about the storage uh spec efficiency together with our underlying like sim provision sna stuff and data matters. So we care about the different data factors here durability

availability consistent a lot of scene and you I will show show all the items in the capacity page later and on top of that we already provide a general function right for the volume percent volume so advanced volume operation encryion extension trim etc all support by longhome and operation finally we come to operation user interface we longhome day one built on the kubernetics So all you familiar

with the Kubernetes user experience operation do the same for long home but we also provide the user interface a simple one that you can quickly understand how to measure life cycle of volume and also in actual functionalities at deployment agnostic we want long can a role anywhere so you can play the local age cloud any place you want to do and virtual action and nowadays virtual action

uh Something change workload matters nowadays together with the containers. Lome provider by default rewrite many volume for migration volume across a different node together with your VN or together this pro together with the cuber more feature you can come to check our documentation. So this is general architecture usually you check from the loo documentation but I just want to let you know we have two parts actually

three parts. One is control plane we call it local managers and the other one is a instant manager not show here but majorly you manage your engine replica and downstream disk. So all the stuff will be a unifi interface from the engine to your workload and m point m point in your part for sure but downstream replica communication will be coordinate by engine and of course monitor

about everything on your know about disk. So for the interface CSI CSI spec compatible is a necessary in the Kubernetes storage solution and we also provide a user interface we also have a internal restful API for coordination as well but once again we also provide a long hon man manipulate the CR manipulate CR operate CR you can use a cube cuddle but uh long cuddle is for

your day zero day one day two for troubleshooting preparation you know this kind of operation stuff. So this is capacity. I will I won't go to each one. I just want you to see the highlight parts. This is key area I want to invest over time from the volume type SS modes and CSI compatible capacity all the things they are in long nowadays. And of course I

mentioned something about the V2 especially the engine migration. This is something we are doing right now. for the future uh lab violin lab aguate for B2 so but I will share more later and overview the longhaul so right now it's a project status uh keep growing uh very stable and if you see we have a mattress is a public mattress matricl long.io You can check this over

twound thousand nodes rhino grows of year over year uh 35% and cluster also uh more than 50% and preform agnosticity is very important trigger point to that the user adapt the longhaul so like I say you can deploy anywhere but I want to emphasize it actually in the different domain adaption as well and also we support different type of operation system including immutable operation system. Maybe most

of some of you use a talos. Uh we work with a taros as well. Yes, this uh just a celebration. This is great achievement at least for long community. We have a lot of installation >> Sorry. And I know some of you may be interesting about how I can contribute longhome especially longhome right now is getting stable getting largely adapted around the world. So we basically leverage

a GitHub project. We run a spring base two weeks and there are some area one is a community we respect for the report from the community user. So is a community spring and we also have a long spring is for the feature developments and Q spring is for the automation stuff. So all the tiki is opens you can check everything what's going on and which ticket in

which status and you have opinion you want to jump into anytime. will come and spring also come with a spring release. So every by end of uh spring we will have a spring release. The purpose is you want people try some we already implement in that spring but also want to get the early feedback and release cadency is resin we introduced in resin uh release probably start

from 1 or 19 right now it's a consistent and settle so basically you will get the minor risk in January May and September and each minor risk we will have a two months patch release in six months actively support. So this is currently happening. We will keep this pattern. Uh yes this about our project and recycle right now it's data engine. I show this page a few

times. This our database. So I just want to let you know right now we have a V1. It's a very in-house solution. we build from the iscasi front ends and database internal TCP server all stuff for the IO and however right now we are building another one for high high performance uh data engine it's based on SBDK and we've adapt the different front end M&M me of

over fabrics and Ubro and also SPDK stuff target servers and down to the logical fing we are using right now the point here is the design is the Okay. So if you're familiar with the V1 you will have a same ideas what's happening what we design for V2. The architecture is similar. So this is a key point for you. Uh for V2 we use the data plan

we come with another component we say uh longhaul SPK engine compared to a longhaul engine inside a long instant manager. If you don't know know about that you just think about is a data plan for long hole and storage firmware we based on the SPDK and front end uh for sure uh SPK is a poly modes you have you need have a dedicated resource for it dedicate

number of CPU but people come with a come with uh to ask if we have a low low spec uh spec environment how I can I use the v2 yes this something we are working on right now interruption mode we want to make the user you don't need to spec just by the number of CPU for the V2 engine. You just use the share resource now like

V1. So right now the implementation is in the middle. We ship uh partial works already in the current version but eventually want to go there. So this is a very key point to that longhorn can adapt B2 can adapt it anywhere and BDF is a framework stuff in the SPK that is easy to extend what we want to achieve for the uh storage level and lo is

a key point to let us achieve the same provision and uh the the snure stuff of course our backup bas based on the snure as well. So again what I talking about here is the same as V1 all the design is the same. This is a major message I want to deliver. If you use start using a V2 don't get um nothing complicated because everything just the

same and front end we also come with a default MVME over fabric is actually TCP but we come with a Ubro as well and we provide some tuning parameters for Ubro and in the corresponding K key if you you you want to know know about the performance there are some information okay so uh B2 performance update This time we do different previously run on the equnis battle

but this time we run on oracle uh OCI and we choose the IO intensive uh instance for our testing. So if you check the baseline the the performance of disk is very good and we based on that to do our uh the testing we usually do the testing for every release. Yeah. Next. So this IO performance uh we do the replica on the same nodes across node

engine cross replica on different nodes and three replics and along with the number of CPU codes you will see the significant increase for the performance. So ray one is v1 and others v2. So you can simply understand why we are doing v2 right now. Uh yeah so we we do the IOPS latency through so all the information is here. I won't dive into but if you have

question you can come later I can but this one I want to special explain you will see why V2 is not outstanding compared to V1 if we use the last CPU because instance manager we we are using right now is a share mode for V1 in our test environment is very focused on the one world load testing so we cannot control very well for resource for the

V1 so this is something you will see that some bias here But if in the real production thing will be changed. Yeah. So this also uh su uh this is let stuff. Yes. Can never change. Okay. Good. So uh v1 um yeah we just come with the 1.11 release in January. So I want to brief uh let you guys know what we deliver. So in 111 we

call it out. Uh V2 is the tech technical preview. The purpose for it we want to see more adoption because we have matrix. We want to know how people use the V2 nowadays. And I also our team also build a dog fooding environment in our QA testing every day. Not just for longhome. Our other project run everything on the longhome. We use a V2. So we run

the testing every day and so there are two criteria to we want we can say okay it's a ready for tech technical review the first one is we catch up with the V1 features and there's a link you can check all the feature about the V1 have right has right now we catch up with that a few we still follow but it will be happen in 112

and increase the test coverage for V1 uh for V2 based on V1 test case so this This current test case we have over 600 test case end to end it's not integration testing or or unit test case this end to end test case including the native test case like involuntary no down clust up this kind of or stress testing for IO and you will see the first

one the in the old test case we implement in pi test we have uh like uh 378 test case right now already achieved 62% can be compatible for v2 too we run every day and the new test case we implement in the robot framework and right now it's probably still over yeah 330 test case all for v1 and v2 so they can run all togethers even we

have more test case view for v2 so this is two criteria we want to say yes he's ready he's a preview ready so you can try >> yeah next yes >> next one >> okay key features The first one suppose uh the ubra fun it's actually ubra fun for v2 is start on one n but he cannot tune the performance. So in 111 you can adjust the

queue size n number of queue you can improve your IO performance depends on your configuration and the second one is a faster replica rebuilding. In the past the replica rebuilding all always come from one healthy replica but right now you can imparel. So we also have a benchmark. You can check the corresponding tickets in the release note. And that's one. Next one is a balancer where uh

the uh replica scheduling. In the past we have a very high priority condition just space available disk space available. But we did not handle well for the concurrent situation. If there are a lot of replica rebuilding happen at the same time, how we can choose the right choice? So right now we including the different factor to cons consider this one to make sure the replica can be

balanced. Well, >> sorry. >> Okay. Yeah, this one uh storage cluster uh allow topology uh if you join the yesterday sessions from our friends never and they talk about the CSI storage capacity great contribution and that the long haul can be uh where together you workload and this one is about troubleshooting because sometimes people will say ah what my my volume my disc have some problems what's

going on so we have a native support for monitoring for V1 disk and V2 and V2 information is from the SPK logical volume for sure but V1 we will have active monitoring per disk and share major networking uh maybe some of you know about storage network we you can set up a uh dedicated storage network for your data plan but rewrite many actually we run the uh

NFS Ganesa for each volume so from our perspective is another storage stack. So sometimes user want to do a different storage network for different stack. So in this release we have a you can specify another storage network for the share manager networking decoupled from the underlying broad device uh storage network and the last one is also contribute from the people outside from sus. So he contribute the

another access more rewrite one's parts to make sure he can ving can only access to the inclusive parts. uh 1.12 V2 is going to G and we are get ready and we are wrapping up to catch up the missing coverage in the uh automation testing and also some uh feature uh we want to pair with the V1 feature parity okay and live upgrades uh when you're doing

the long upgrade the long engine will do a live upgrade this one will be the functionality will be implement into 112 but start to support from 112 to 130 the behavior will be a little change uh different from v1 this is only different because we need to deal well with the spk demon but more information we'll share later not later in near future so IPv6 IPv6 uh

we already support for v1 uh this special for data center stuff and we were supposed to in B2 and even we have a hyper mode so but mix stack but this is next story and engine migration if you know Longhorn support the lab upgrade long support the volume lab migration for the VN stuff right now we are going to support the engine migration so you mean you

can your client initiator can run no A but your run other nodes so this is a foundation and preparation for our V2 life upgrade but it will be a individual uh standalone feature as well and open stuff uh often we already for V1 so we want to do the same to make sure you can do the leftover resource clean up this kind of thing automation stuff the

same experience as V1 fast cone fone is a speed up the volume creation because we want to le leverage this nation we want to wait until all replica rebuilding. So we want to if you have a is in replica from source volume we want to just bring up. So this is the first one we want to do for V2 but V1 will be later we will see

how the feedback first and new features uh in place resource resize uh feature from the Kubernetes 13 uh 33 we want to do for the instance manager because random instant manager will be you need to specify the request uh the the the compute resource but this one will be changed and offline replica rebuild building for the data protection and the volume offline. How we can make it

happen the rebuilding automatically. We already have that behavior but this setting is forced. You need to enable but we want to make it enabled by default and rebuild uh replica review at runtime also best efficiency. Uh the last one uh sharding uh the sharding we will come we will based on the spa uh the BDF design come out with the erasial uh erasio coding uh BDF on

top of it. So this is time to unleash the volume size uh the capacity from the longhorn because in the past longhorn only support not past currently long only support replication mode. So each replica will be the same size of uh the replicas across the different n and and disk. So sharding can unleash this one and sh data priority all togethers. So now. We will come with

the one uh we call it long enhancement proposal there will be a PR soon. So if interesting about the detail you can jump into yeah there's one. Okay very fast. So, so if you have any question you can come there community update. >> Yeah. So, uh going to run by quickly because I realize we have like 10 minutes left and I do want to give you all

an opportunity to ask David questions. Uh but this is to say that we are very much planning to move towards the next level of um the CNCF ecosystem which is graduation in the coming year hopefully with everything that we have planned in terms of technical updates. However, uh towards that we need your help which is why I come in. Uh so current status um and uh future

road map is what I'm going to be discussing fairly quickly. The numbers um I'm not sure if you all are interested. However, I think it's um it's an overall health picture of the current community that we have along with all the technical updates that David has given from the GitHub side of things. However, what what I'd like to highlight is that we are making good progress towards

engaging nonsusa contributors. Now, why is this important is because we are trying to um you know build out a board of maintainers really that can take this project um ahead into its next level. Uh and as much as uh we at Souza love the project, we also want for you all as the end users and as um you know community members to contribute to it and have

a say in the way the project proceeds. So um towards that we've been uh having these avenues opened up in our uh GitHub for uh welcoming code contributions and other ways. But um to the point of the last release uh that is version 1.11.0 or we had around actually not it this is actually incorrect because I've written a tilt here but it's it's actually greater than 50%

uh they were non-sou contributors there were people who reported bugs uh opened up issues obviously but there were also people who contributed actual code to the version that y'all used in production and I know that this seems like a very minor thing but it is extremely important to uh continue with the trend of vendor neutrality and to ensure that the Longhorn project is not just spearheaded by

our company. So uh the most popular modes of contribution were uh are listed here and you can see the number of people who have contributed to this release overall and um one of the things I'm not going to play the entire video because really that's not relevant. You can actually check the video out. David has dived deep or I'm I'm sorry for the English but like the

there has gone more in depth uh with respect to the Longhorn version 1.11.0 in this global online meetup that we hosted last month. So if you're more interested in the features you can just go to the slide deck after the event and you'll be able to view this. Um another thing that uh we have been sorry this is going to start playing and another thing that we

have started post um I think June 2025 is hosting monthly community meetings. Again, this is in an effort to bring more of you into the fold in terms of feedback, in terms of contributions, not just in the code side of things, but also really documentation, bug reports, etc. So, we alternate between APAC and um American/EU friendly times. uh time uh zones are listed and the times are

listed here but um recordings are available on YouTube for each of these meetings. So these are recorded so you can view them on our official YouTube channel and this is of course if the meetings are in quorum because there are instances when people do not show up so then you won't have a recording so recordings are there uh and also if you would like to be a

part of upcoming meetings um thank you for the notification um if you would like to be a part of upcoming meetings uh please do check out the agenda do the zoom link the CNCF mailing list everything's available there and uh we highly uh recommend that you join us because that's a place where um other maintainers not just David and myself but people from other parts of the

world do uh come in and actually walk you all through the current status of Longhorn and if you have doubts you can actually even ask us there. Um, next one is um if yall are not really um uh looking out to engage with us on Slack, we understand if you just want to stay updated, we've opened up these new avenues for contribution. Not asking y'all in a

typical influencer way to like uh like, share and subscribe. These are just avenues. If you'all want to stay updated so along with the GitHub, if this is something that you'd like to do, please do so. Uh we will have um we will be doing the normal uh CubeCon um you know talks and everything else but uh given the flurry of updates it's difficult so we highly recommend

you to stay updated by following us on social media but more importantly is the adopters section. Now, I'm not sure if any of yall were there during the TOC uh AMA that was hosted in the project pavilion a couple of days back, but one of the things that they actively mentioned there is that uh the CNCF TOC has yours on the ground and they actively listen to

the signals from the end user community towards assessing if a project is actually um performing well, is actually solving a problem. And the only way that that can be done is through sharing your stories um and showing your support for the project that is there. So if if any of you uh here use long on in production uh please do consider um having yourself entered into the

adopters file because it's extremely important for us when moving levels to um sort of you know justify to the TOC why this needs to be moving levels except the technical compos competency in itself because you can write beautiful code uh but if it's not actually solving a problem it's not really useful to anyone. So if if we are helping or if the project is helping you out

in your production use case, we request you to please have this um adopter file entry in your mind when you go back to your day jobs and u uh have you know yourself entered in there. uh but that's to say that uh the future roadmap for the next year when we are embarking on this move to graduation is to improve project health not just from a technical

perspective because yes that's important but also to broaden uh the areas in which you all can come in and contribute uh firstly that is looking very um uh bright right now because we're building out our governance structure uh and this is important I I know that uh as engineers a lot of people are not interested in the governance structure. You'll just want the project to come here

and just function in your you know environments. But uh to ensure that the Longhorn project does continue um we need to build out a governance structure that's representative of the community that it serves. So we are identifying um and dividing the project currently into areas so that we can have appropriate roles staff for contribution and we are going to invite community folks to come in and staff

some of those roles. It's not going to be all soua folks. Uh we have a wholesome community right now with uh consistent contributors. So we do we're currently in the process of identifying and dividing the project. But outside of that, we also uh want to be transparent about all of this. So we are going to be um ensuring that the roles within the project are documented and

we have a community ladder to show the ascension and we want to broaden the pool that we have currently because 72 people across the globe is not really a great start. Um it's good but it's not a great start. So, we want more of you all to come and join us. And uh we want to be growing y'all into maintainership as well. I mean, if you'all are

interested because this uh as much as we love the project, like I said before, we want this to be representative of the community that it serves. So, if any of this interests you, uh let's talk not necessarily today, not necessarily in the next week, but y'all can come during a call and show your interest if that's that's how it works. And with that I think the stretch

goals because I have been um sort of molded into this role to also provide my um extension. Uh so the stretch goals um are to participate in ecosystemwide uh programs and to ensure we have cross ecosystem collaboration as well with other projects. Uh so if any of this is interesting and if you are interested in talking to us about any of what you saw in the past

30 minutes, please do catch us in the hallway. We are over time and I'm really sorry about that. But uh y'all are yall are a lovely audience and thank you so much for attending today.