TAG Operational Resilience: Sustainability Month To... Mario F, Alolita S, Carol V, N Pal & Saiyam P
About this talk
This session discusses the updates and initiatives of the TAG Operational Resilience group within the CNCF ecosystem. The speaker, Alolita Sharma, elaborates on the restructuring of technical advisory groups and the creation of focused areas such as observability, business continuity, and resource optimization. The group's mission is to establish standards and practices for resilient, observable cloud-native systems that operate efficiently in production environments. Several initiatives are highlighted, including the development guidelines for project releases, the creation of a reference framework for service reliability, and efforts to identify observability personas. The discussion emphasizes the importance of community participation in these initiatives and the need for collaboration to enhance operational resilience across various projects.
Full transcript
Hi everyone. Good afternoon. So we are uh part of the TAG operational resilience uh group and we have our chairs as well as our technical leads presenting updates on the different uh efforts that have been ongoing at the TAG and also talk a little bit about um what's happening you know what has happened so far and where we are going next. I'm Alolita Sharma and I also
have our esteemed chairs. Uh Rafa is not here today but Mario's here. This is Mario and Sam is here. So these are our chairs for the tag. And then we also have some of our technical leads here. Uh Nabun, >> hi. >> Um Carol, and then we don't have Matt here today. I'm Alita. And uh I think Rafa is not joining us either. Right. Okay. So with
that said uh let's move on again uh just wanted to reiterate that one of the things we do follow on the tag all the meetings as well as the discussions that are ongoing that please treat each other with uh kindness and respect and that goes we know what goes around comes around. So it really matters a lot and of course uh please look at the CNCF code
of conduct if you haven't looked at it before. Uh it again reiterates what we just said. So I want to talk a little bit about you know the um reset process that was uh put into place by the TOC for the CNCF uh last year. It was a reboot of the previous tag structure in case you guys you know uh were not familiar. And fundamentally all the
tags that existed in the past before last year were re um grouped and what that did is that that created you know four five specific areas of work with tag operational resilience as one of the key areas. We'll talk about the mission but you can see like areas such as observability which was some one of the areas I was working on uh management business continuity resource optimization
cost efficiency green initiatives with energy performance troubleshooting reliability and dayto ops were all kind of converged into this broad area. Um similarly there's tag developer experience, tag workloads foundation as well as tag infrastructure which are the two or three other tags which are horizontals and then you have another major vertical which is the tag security group which was also you know focused on all things security. So
just wanted to kind of reiterate that because that's how the tags have gotten restructured. But the discussions around operational resilience and what it means to be production ready have all kind of converged into this operational resilience tag. So moving on uh again I'd like to kind of reiterate what our charter is and then I'll hand it off to Darun to talk more about the initiatives is that
as a mission what we you know kind of uh discussed with the TOC and then evolved upon was to define practices and standards for building operating uh adopting and managing resilient observable and efficient cloudnative systems. It's a lot, right? Because this actually spans pretty much, you know, what you have to do in order to build very large scale cloudnative systems and services and actually run them and
operate them in an production ready way, right? Um, and fundamentally it looks at systems, services, applications, architectures beyond initial deployment. So once you're done development you know done with development with testing then you go into operational you know running these large services in prod and that's what this tag focuses on and disco discusses. So with that said, I'll hand it over to uh and of course I
cannot forget our beautiful little nice icon. Uh and uh I think Siam was one of the initial proposers of this cool tardy grrops. Um but it has you know all the initials and maybe Sam you want to say a few words about how you came about this cool you know. Yeah, I mean it's it's all uh like obviously I used AI and then some of the other
prompts to generate this but we did get to manage some stickers. Uh I don't think we have any right now but uh but yeah this is something that represents uh resilience. Um and that's all the tag is also about you know operational >> Awesome. So with that said again we'll kind of hand it hand it off to Nun and Narun is going to talk about all the
initiatives. >> Yeah. So we kind of have to like work towards some goal um as part of tag operational resilience and the work is all organized as initiatives. So they are basically light uh organizational units you can say used for any kind of TOC or tag work and they need to be like very scoped and time bound. So they can't run for like really long time they
have to wrap up with some defined goal. Um we can talk about now all the initiatives that tag O is leading or co-leading and running. Um moving on let's see what's the first one we have. we have the project release guidelines um initiative. what do we aim to do? So what we figured out in the past few months is that many of the CNCF projects which are
in sandbox incubating or even a few graduated projects they wanted more guidance on how to formalize their release processes or structures on how do they manage code, how do they manage artifacts, how do they publish things, how do they distribute their artifacts, where do they distribute, how do they ensure supply chain security. So the idea of this initiative is to create guidelines, show them patterns and provide
them with a few reference implementations through examples of existing projects to help them secure make their workflows robust and repeatable so that they can sustain. Now one key action item that um we want all the people here and whoever would be seeing on the recording to do is please scan this QR code. you will reach the draft guidelines uh white paper. We really want you all to
come and comment on that white paper and if you want to participate as a project maintainer who wants to give us some more thoughts or requirements on what do you actually need from us come and comment on the doc or reach reach us reach us out on either the tag operational resilience channel on CNCF Slack or there is an initiative channel as well called initiative project release
guidelines. So we watch those spaces specifically for discussions related to project release guidelines. If you want to help out in any form or fashion or in terms of like giving us more requirements or giving us feedback on the white paper or even like participating as part of a case study um if you as a project maintainer want to impart your thoughts on what you do. Um with
that I hand it over to Carol. >> Yeah. Well, here is another initiative that is running under the TAC operational resilience that is the initiative reference framework for the levels of service reliability optimmentation. That's very long name but it's as you can see this is a good sample how it started like one month ago launched by Severing maybe you know severing from observability. So uh the inspirations
was try to create uh uh levels as you can see this this was in driving automation inspirations but we were we would like to transform this in a white paper about service automations and how is working uh someone's has an idea that is related to our t and you can open this initiative that was severing that he join us to the tag um we have the slack
group that will be very important And if you are interested to to be part of this white paper or participate is just started two one month ago like two meetings because it's be weekly and yeah I think that will be my overview about this this this initiative. >> Okay so uh thanks Carol. I just wanted to quickly go over an initiative that is actually open for contribution
and uh this is a newly u you know proposed initiative on identifying observability personas and why this matters is because you really have a very large diversity of you know different types of users in you know how observability is used right that is you could have uh operators you know on one end, but you also have developers, you also have other, you know, consumers of observability. And
so this initiative is really, you know, kind of has the following goals where we are kind of working together as a group to identify user personas uh based on actually talking to different uh end user groups uh and companies and also building a common shared definition for the entire CNCF you know project ecosystem. And why that matters is that you know again there are different parts in
operational resilience components right and uh ser observability data telemetry and metrics traces logs profiles are all consumed in different ways by these observability personas. So kind of identifying that relationship and having a clear reference for all the different CNCF projects to be able to use. So when they you know say talk about a developer what does that mean right and and what do they use primarily versus
second as secondary signals and similarly uh this also helps end users in being able to take you know observability definitions which they can use in their internal strategies across their own organizations as well as organize you know what the features they are building for different types of personas. And last but not least, you know, kind of reduce the alert fatigue so that everybody's kind of on the
same page when you're referring to different personas because observability in itself is also evolving as all of you know, you know, and you probably I'd like a show of hands here. Uh who all are thinking about AI observability now, right? Everybody. So, so that said, uh again would really encourage folks to go and take a look at this issue. the link is right here and learn about
the details of what we are trying to do. It's a it's a great way to get started and uh really dive into you know this project and if you have any questions again please reach out to uh any of us or me or Matt who are kind of working on this initiative or please also join the tag operational resilience slack channel. All right with that said I'll
hand it over to Mario. Yeah. So the cloud this is an initiative that um which is basically not accepted. This means that um as Carol briefly touched it, anyone can come up with an initiative. The only rule that we basically set by ourselves is that we don't accept initiatives where we don't see traction in the beginning. So basically this is something that people can open an initiative
but it will only start once we see like a couple of commitment to a specific topic and I think that's something that uh what we are still struggling with or what is still um hard for us is like to communicate uh to the broader folks to say like hey we want input and we need help to run all of this stuff because you see we are like
three four so We are n eight people who are like organizing and tech leading stuff but obviously we can't run all of those by ourselves and we also don't want to run all all of those by ourselves because we want to get input from every uh from everyone and uh this is a good example. So the idea is that we this is aiming to publish a white
paper to say like how do we create um the world that it's it ensures business continuity from using open source projects and how we move on from there and basically say like I want this is what I need to do from a project side to ensure that companies are using my product and are continuing to use my product and basically have like also like what is an
incident happening, what is happening with uh if I don't have backups and uh also on a project level and on a global level. So if you're interested on those initiatives and even if they have not started um please just put your name into the GitHub issue to say like hey that's something that I want join um so that we then can say like okay if we have
like a critical mass of people that is actually working on this that we can work on such in >> Yes. And that goes for all of all of the initiatives we're calling out here. So uh with with all the AI things that are happening, one of the key things is uh sustainability and in that um whether you talk about the rising of the AI factories, NeoClouds uh
whatever is happening more GPUs, I mean Jensen is creating the racks which will take more gaw of uh power to power up the GPUs to create even better models and even better things uh for the people to consume but in the end uh the amount of energy that we have is limited which is where we see a lot of innovation happening in the sustainability front. Uh with
that there are some of the uh things that are part of the tag operational resilience. So anything related to sustainability falls under the tag. Uh there are certain initiatives and certain things that we do. So yearly we organize uh the cloudnative sustainability week. Now this is a pure community-led like the tag itself is all community-led but uh this initiative is uh like people from different parts of
the globe come together and they organize these local meetups and these can be in person virtual and we have been doing it for a few years now and um as part of restructuring which happened last year we still managed to kind of do it um so these are some of the pictures from Tokyo and Barcelona Um so cloudnative sustainability month is something that is celebrated in order
to tell people about you know the green software foundation SCI some of the uh projects within the CNCF ecosystem that you can use how to efficiently utilize the resources because in the end how you implement uh sustainability at your organization you need to tie it with cost. So the more you utilize your infrastructure the lesser you will waste the resources and indirectly you are doing sustainability for
your organization and in turn you are saving the costs as well. Um so so that's how the that's how the whole picture is tied and these um events are very good to get into locally with the local community uh to understand about and we have representatives from different parts of the globe who drives these um so this happens every year it'll happen this year as well uh
again uh the tag operational resilience CNCF of Slack channel is the place where you should be uh because that is where every communication every detail is posted. This is also posted on CNCF as a blog uh that we are organizing this. So also keep an eye on CNCF blog as well. Uh another project uh under the tag operational resilience is project green reviews. A very interesting project.
It is uh basically to uh review uh the provide the metrics and the guidelines for measuring the sustainability footprint of CNCF projects. So that is the goal. Uh the project needs help um because it needs more contributors like it needs serious help because this is a very good project. Uh and if this if this works out for all the CNCF projects, it'll be really good to understand
their carbon footprint. Um and Kepler which is also one of the CNCF projects um have been you know kind of involved in this particular domain. And then uh there are people like we have uh meetings both in the APAC time zone and in in the uh North America time zone. So we have two meetings that that is done for the operational resilience. So you can join any
one of them with respect to the time zone that you are comfortable with. Um this is where I believe it's like CNCF is all about collaboration and you know keep moving forward is the theme for this particular cubecon and cloud native con. And my request would be that this project is really very interesting. If anybody would like to kind of contribute to code and is interested in
sustainability um getting their hands dirty, I think this is the place where you should get involved with and because this has very less contributors right now and it really needs more uh more people to contribute to get this over the finish line. Like Mario said, um any initiative that we kind of put in, we have an end goal and we we want to put only those initiatives
that have kind of an end goal where we know this will be the end state of a particular initiative once we achieve it. So we don't want to leave that in a limbo zone. Uh we want to achieve something that we started. So >> and also to chime in there for one second. So uh I'm involved also in a lot of project reviews for moving projects. So
when projects are moving levels like from sandbox to incubation and stuff like this. So they we have the security review, we have the governance review, we have the uh technical review that you need that needs to be fulfilled and it is a mid to long-term goal that the green review also becomes one of those requirements that need to be filled out for every project. Um so that
we have like this um yeah the also the sustainability side in in this whole review process and this is then also published uh in the TOC repo when projects are moving levels. So this is like really something that that would impact the whole ecosystem in the mid to long term. >> Yeah. and and Nikki is also here like one of the leads for uh the green reviews
project. So if you have any questions, she's there back. So if you have any questions, you can reach out to her as well. Um yeah, it's it's a pretty good project that is you you would love to get involved with and you can do create an impact. >> Okay, Tom. Uh let's see with the this slide that in theory everything is easy. how creating initiative that you
have to go to the TOC in the CNCF GitHub project and you try to do an open issue and you will see their initiative but I think yeah you have to fill up some some some spaces but I think the most important is go to the meetings that we have these meetings in D operational resilient and talk with us if you have some idea because that will
be uh the best how to start it and also uh as um Mario said we we hope that the initiative is no more than three months no more than three months that we can accomplish this white papers or any initiative that you have it. So I think uh good advice is have at least minimal three person that are interested that is the the the case with the
other initiative business community that is only one person interested and I think unless quarum of three and we can help you I don't know by social networks or slack to try to get more people but unless you try to have a a pyramid I don't know three people I think of your trust that you can follow and push this And yeah I think with that h also
all these ideas have to be volunteer on this don't you have all these ideas inside our tag that will be observability business continuity that already we have a stop issue there uh resource optimization if you have any idea about these topics you can uh come with us and we can help you to develop the idea and try to support you that's is will be our our job
here talk with us Mario. Yeah. So, basically we have bi-weekly meetings. Um, we have them one is Apec friendly. Uh, which is usually run by SIM in Nun because time zones. Uh, we have then the US friendly one which is uh run by other folks. Yeah, it's so so yeah EU friendly. And um so come in those meetings. We usually publish the agenda. You're also free to
put in stuff into the agenda that you want to discuss. We have our Slack channel and uh yeah we are looking for input from everyone. Um you can do an impact and this is mostly non-code relevant input. So this is also something who is not like hey I'm a programmer and I don't know I I other program but I still want to contribute to open source. That's
basically a good way uh because we are publishing a lot of white papers and structure for the cloud native community and with this thanks everyone and questions. >> Yeah. Do we have time for questions? Yeah, I think >> we have >> Yeah, please come up to the mic >> working. Yeah. First of all uh before coming to uh question thanks for addressing energy first then the cost
because actually when you lower the energy as you said uh which in this cubecon at least I have seen one poster from Henry and five six sessions which was not happening in London or Paris. So I am very grateful that energy is in the first place uh in that sense. My questions is uh one uh observability itself is a big topic and it's uh uh developer experience
network wherever there are also many uh other than observability vendors solutions many products uh showing it. So it's a when you put it on one uh silo you automatically uh push for example de developer experience without observability but also they have a lot of observability there. So uh one part of the question is about uh I'm not judging the decision but how you concluded to group it
uh under there is my question. I want to learn how did you reach there? uh and uh also related with it in tag observability per area is here almost a three years work great work has been delivered now it's mentioned 3 months uh so yeah that 3 months is a little bit scary actually that's all thank you >> so um when the when the restructuring happened basically
the t mostly the TOC sat down and the the tab so the the end user group sat down and basically really figured out what would be good topics that fit together and um there was always the or there's always the the idea of tags are not exclusive by themselves. So there is like the um the idea is always that there's a cross collaboration within the tech. So
for example I am involved in two initiatives who are in tech security because they are touch they they touch the topic of operational resilience because it's supply chain but supply chain is mainly under security because it's a security topic. So there it's not like a we exclude topics from each other. we just look where does the core fit and then there's a cross tech collaboration. It's mo
mostly it's it's just an organizational layer. So uh someone someone does like calendar invites managing where where it's located but it's not like limiting the topic but you someone needs to be in charge for the yeah for the for the business stuff you know what I mean? Um and that's basically uh why observability was into uh uh operational resilience because they every when projects need to fill
out the graduation projects to go to graduated um it's observability is the main observability questions come in day two and that that's the reason why day two operate it's the form that we have from a technical review and uh that's basically why it was moved into tech operational resilience and to answer the Second question with the initiative limitation of 3 months there can be longer terms because
there can be projects coming out of those initiatives that can continue there but the initiatives is basically to kick off things to start working. So for example the green reviews project will eventually become a sub project of tech operational resilience with no time limit whatsoever. It's basically like to kick off things to and not to create white papers that never get finished because people leave. There was
there were in the old text there were white papers that were started and were written for one and a half year and never published because they never got to the now we publish a state. >> Hi. Hi. Yeah, I have a question about um sort of AI in the operations. Um I was wondering if operational resilience covered any of that. Um just like AI in the workplace
and making sure that people don't destroy their environments >> to cover, huh? We have I mean AI is infecting uh sorry affect >> a affecting and infecting uh all of all of the topics. So it's basically cross the whole infrastructure. What we will eventually do is and that's I think Alita already pointed this out in observability we have like special requirements. So we will pick AI related
topics that are important or that fall into our categories that we uh that we want to look at. Um we cannot generalize or AI into one of the texts because it's basically in every in every part of this. Um I think the whole AI uh yeah I I send a prompt and I there's basically one tree less on on the planet uh thing. Um that's something that
is part of the uh what's Sam presented with the sustainability week and also the green review thing is how much is my project actually burning. So for example, if you you if you take cubeflow as a project, cubeflow needs GPUs to run its tests. That's something that we where we then can measure what's the energy consumption of the project just to run the infrastructure. What are the
implications of your company if you run cubeflow and and so on and so forth. So um this will the topic is broad and it will definitely be part of any tech operational resilience conversation that it's already been happening. >> Thank you. Um uh a comment that I have here I think it's a bit relate to what Carol mentioned and and and Alita um because on on the
consideration for the the res reliability automation levels um there is I'm ignoring the details but like there'll be there's like the with no automation there's like with like some automation there's like with a lot of like fully automated right uh fully autonomous um I think do did we have some consideration the personas Because I think at least the middle ground here there is like the because usually
when define persons is about how you fit CJS right so how we consider the critical user user journeys that those persons will be used for and I think like the assisted the one that like the agent is assisting me that uh that should be I don't know uh on call and on call with agent right and whatever like engineer launch a feature and engineer with agent or
whatever is that part >> in fact if you look at the issue there are different uh examples that are listed but again please feel free to add to the issue because again it's an evolving and you know all of us together kind of really uh can provide feedback there and get you know get that more finely involved because I think even as you highlighted uh the idea
of AI assistance right in each of these categories also is kind of another category now in itself and another persona in itself self really. So whether that's uh with sustainability or with observability both of them are equally influenced by that balance >> and um I would definitely say that we should you know kind of comment on the issue. Uh it can be easily added and of course
the objective is that you know within um 90 days of starting the initiative then we can capture some of these areas and actually clearly define them. And also the other thing I'd like to call out to what Mario was saying is that uh if there are you know initiatives that you guys want to see like specific ones it helps to have you know a short kickoff initially
kind of take and define the scope you know in a focus group and then go after a larger implementation or a reference architecture. >> Thank you. Yeah. And also um when an initiative ends the outcome of an initiative can be we need a new which is then more specialized. So what I could see is like after the um persona initiative finished that we then that then the
new initiative is right uh reference architecture or write like a requirements sheet for each of the different personas. So that's it's it can continue right. It's it's not a now it's done and now we forget it. It's it's always evolving, right? Yeah. Thanks,
More from this event
See all 436 talks →
Best of KubeCon + CloudNativeCon Amsterdam 2026
2:17
The Quiet Work of Forever: Sustaining Open Source Communities - O. Hope Amaechi-Okorie, JSON Schema
26:24
Evolving KServe: The Unified Model Inference Platform for Both Predictive and... F. Spolti & J. Lee
32:40
Preventing S3 Cost Storms: Applying Cortex’s Efficiency Lessons to I/O-Heav... A. Fishman-Lichterman
5:32