KubeCon + CloudNativeCon Europe

Introduction To Tag Infrastructure - Kashif Khan, Ericsson & Dylan Page, Lambda.ai

29:18 · 23 Mar 2026 – 26 Mar 2026 · YouTube

About this talk

This talk covers the structure and evolution of the Technical Advisory Groups (TAGs) within the Cloud Native Computing Foundation (CNCF), focusing specifically on the TAG for Infrastructure. The speakers, Dylan Paige and Kashif, explain how the TAGs play a critical role in overseeing infrastructure-related projects and fostering collaboration across the cloud-native ecosystem. They detail the TAG's mission to define best practices and standards for infrastructure technologies, including compute, storage, networking, and edge computing. The presentation highlights ongoing initiatives such as the storage landscape white paper, AI data storage, and infrastructure lifecycle management. The speakers emphasize the importance of contributor engagement and outline how individuals can participate in defining the future of cloud-native infrastructure.

Full transcript

Hello everyone. My name is Dylan Paige and joining me is Kashif. >> Yes. >> Uh and we're here to talk to you about tag infrastructure. Um I'm a maintainer for the sandbox project Atlantis. Uh I'm also co-chair of the tag infrastructure as well as the Kate sig infra uh and engineering manager for the core infrastructure team at lambda AI >> and myself I'm also one of the

co-chairs of the CNCF technical advisory group for infrastructure which is why we are here and apart from that I am a maintainer of a project uh CNCF incubating project called metalcube.io IO and I'm working as an open-source architect in Ericen software technology. >> Awesome. So, one of the things that happened about oh nine months ago in June is that the uh CNCF uh rebooted the technical advisory

groups. Now, a technical advisory group is the primary organiza one of the primary organizational units under the technical oversight committee. Uh there are 11 people on the technical oversight committee appointed in various ways either through the governing board uh governing body, the end user technical board or from the TOC itself. And so with those 11 people they are responsible for the entire technical vision and uh life

cycle of the 240 projects in counting in the CNCF. So that's that's a lot for 11 people. So one way they uh solve this is by delegating through these technical advisory groups. Um these technical advisory groups oversee and coordinate interests across the CNCF projects and various um initiatives and sub projects and community groups and the broader cloudnative uh community. Uh our purpose is to scale technical contributions

while helping to maintain quality and advancing the CNCF's mission of making cloudnative computing ubiquitous. We serve as a bridge between the CNCF projects within our domains. We provide technical expertise during project reviews and evaluations. We identify gaps in the project portfolio. uh and we also uh foster project maturity as well as educating users with unbiased and practical information. So I mentioned the tag reboot prior uh over

the past 10 years there's been uh a growing number of tags. There was up to eight uh and there was a need to kind of consolidate and refresh the advisory groups. um the uh the ecosystem has been changing and growing over the past 10 years and so this is was done to better address uh the ecosystems needs um one of the ways this was done is we

consolidated the tags from eight down to five uh to allow for better overlap with the TOC liaison as I mentioned there's 11 members five tags means each tag gets two TOC liaison yay human load balancing um working groups were too ambiguous. You know, we were dealing with various issues of um people would start a working group or they would drag on and scope was not clearly defined

in the beginning and it was hard to really kind of nurture and help these working groups uh along. So those have been replaced now with sub projects and initiatives. Initiatives are now short-lived, tightly scoped and time bound. While sub projects are ongoing or permanent projects or programs that require long-term stewardship. Um some examples of these are like the artificial intelligence initiative uh as well as the project

reviews sub project and contributor contributor strategy sub projects. Now these report directly into the TOC uh and representatives from the tags do participate in those but it's not limited to just the tags. Anyone can join these initiatives. Anyone can join these sub projects and these and there are initiatives and sub projects that are also uh under the tags perview as well not just the TOC. Now here

are the five tags we have tag operational resilience developer experience security and compliance workload foundation and infrastructure. Um you can see the way this is kind of laid out there are three main t technical tags. Well, they are all technical but as you can see there's the intention here is there's the layers uh starting at the developer experience at the top working through the found workload foundations

orchestration all the way down into the infrastructure layer. Uh this is purposely intentional um to promote easier collaboration and overlap between the tags. As you see, we also have operational resilience and security and compliance being uh adjacent vertical tags because these topics are uh conducive to all of the three uh layers. You know, security may mean different things for someone as a developer versus someone who's working

with infrastructure. Same for operational resilience, observability, uh performance, and etc. Here are some QR codes that point to the historical GitHub issue for the tag reboot on the TOC as well as some presentation slides uh that were uh presented uh uh last year and over the summer for folks if you want his deep deeper historical context on the change. All right, so specifically TAG infrastructure, that's what

we're here for. Um who are we? Um we are uh we we are made up of uh there's seven of us. There's three chairs and four technical leads uh each serving either one to twoyear terms. Uh and we work with and report to the wonderful Ricardo and Karina as our TOC liaison. Um we primarily focus on storage network data DNS compute service mesh infrastructure as code load

balancing edge and sovereignty. Um, and we also have a booth in the project pavilion P7A. It's in the back of the solution showcase uh during the mornings. Um I also just wanted to point out that uh the uh Kashif will be going in deeper about our charter and our mission. But one of the purposes uh of the tag is uh is that it's a living group and

so that these um these are the primary focuses that we decided on when we uh rebooted the tag, but these can change over time to address the needs of the ecosystem if needed. >> Yeah. So uh I'm going to do a kind of a deep dive on the charter itself and as Dylan said that it's a living document. Uh it's it's should evolve uh if we identify

gaps from the contributors. But anyhow the current version is it's like uh says that what is tag infrastructure doing? So it's it is kind of sitting at the very foundation uh of the cloud native computing foundations technical landscape. The mission is to define and advance the best practices uh standards and assessment models around the infrastructure that powers the cloud native systems. Uh where other tags as we

saw in the last uh slide that uh they might focus on workload observability or application delivery. Tag infrastructure is focused under the hood the compute storage network and the control layers that everything else uh depends upon. The key word as you can see from the charter is that uh it's scalable, resilient, secure and performant. Uh which is where we are trying to set the benchmark uh for

what infrastructure in a cloudnative world should look like. Uh the tag infrastructure is existing mainly to help adapters and to to help adapters solve the real infrastructure pain points like portability across clouds or um consistency between age and and data center environments uh performance trade-offs for example compliance and automation and and and those sort uh sort of things. So it is uh again explicitly aligned with the

CNCF technical oversight committee's uh charter meaning that we are not working in silos. The per the purpose of of tech infrastructure is to keep the entire CNCF uh ecosystem technically coherent at the infrastructure level. So in short um tag infrastructure is trying to define uh what cloudnative infrastructure should look like what principles are we trying to adopt uh in the ecosystem or um what trade-offs are we

making and how we are ensuring that these building blocks are uh uh they do remain open composable as well as interoperable. Um Dylan shortly mentioned about the core technical domains that uh this tag infrastructure is is uh trying to look into which are kind of divided into six. So these are the main building blocks which I'm going to again show some example projects that fall in between

uh these areas each. So data for example as you can see it it refers to all the digital information that is created handled and utilized by the applications. And when I say data, it's it can be both structured and unstructured forms of data. Um it's not limited to u but includes system logs or uh configurations and performance uh matrices as well as uh permanent uh records like

database entries. Uh when we are talking about data in the infrastructure um it's not only about u the matrix it's also about the telemetry uh u uh part of it and as well as the user generated data that can drive the feedback loops as well. We already know projects like fluentd which is kind trying to focus on let's say log routing on transformation and delivery grant tries

uh while um we all are mostly familiar with promethuse and how it actually handles the matrix collection for example for us tag infrastructure we are trying to look into the full data life cycle. So from generation to ingestion, transport and retention with emphasis on schema consistency, back pressure control and cost of observability. So the goal here is again for us is to define data patterns that are

scalable, privacy aware and cloud agnostic. When it comes to storage, uh it can span a wide range of systems as we all know. But um as you can see also from the uh charter that it includes the block storage, file systems, object stores, database, key value stores, messaging platforms and caching layers. So rather than uh focusing on specific technologies, the key is here to understand the fundamental

characteristics that guide how we choose and use the storage for different workloads. Um and then how do you actually evaluate the storage systems because we all know that there are things like availability is very critical because systems must must remain uh accessible during the failures. Uh durability is is is uh critical to ensure the data is not lost even in if there are any hardware drift or

system failures. Um so the cloud native storage is is is about balancing these characteristics to meet the specific needs for each workload. Um some example projects here like Rook which orchestrates a distribute uh distributed backends like like SE uh CubeFS uh explores the file system abstraction uh optimized for Kubernetes uh and then of course uh Longhorn is is trying to demonstrate the lightweight um cloudnative block storage.

So the idea or the aim for tag infrastructure here is interoperability that any workload that uh should understand what a gold storage class means across vendors. Well networking is one of the crucial uh if not the most crucial part of the cloud native ecosystem. It's all about uh enabling as we know the reliability in connectivity and of course with um like int intelligent traffic management across uh

the distributed environments. At the foundation as we know we have these core building blocks like DNS gateway load balancing and service discovery. On top of that sits the service mesh which provides more advanced control over the um uh interervice communication. for example, including routing uh retries and and whatnot. And then the policies enforcement in this layer is also critical because it helps enforce the security, access control

and compliance requirement at the network level. So some example projects again selium for example as we know uses ebpf to provide programmable data paths and fine grain visibility. STO um uh is also apart from other features it it's also uh trying enables us to add policies for resilience uh through service mesh control uh envoy has uh the L7 proxy layer for example so again the idea for

the tag infrastructure is to have some sort of a well- definfined network behavior uh contracts that can make cross mesh and cross cloud uh uh routing predictable Compute is as we know is the execution substrate from the kernel boundary to the specialized hardware and in the cloud native environment especially as we have seen now that when we talk about compute it's not only limited to CPUs we

now need primitives support for the specialized hardware as well for for like GPUs TPUs FPGAAS and whatnot so and especially for high performance and AIdriven workloads. So the the foundation uh of of compute is is to actually enable uh everything else in the cloud native system. So combining the isolation and support for increasingly diverse and specialized hardware also um I have given some of the example projects

as we know like container uh d or cryo which implements the OCI runtime interfaces but then also cloud is is is gaining a lot of traction um in the recent years which extends the model with lightweight capability based web assembly runtime as we know. So u the idea or the goal of tag infrastructure here is to provide guidance for runtime selection or accelerator awareness and sandboxing best

practices. And of course uh goal is again to have a consistent compute abstraction whether the workload is running in a container on virtual machines or bare metals or in wasome modules. Infrastructure management is is in a cloudnative world is about controlling the full life cycle of the system across public, private or on-prem as well as edge and hybrid environments. So at the core we rely on infrastructure

as a code both declarative and imperative approaches. Um so for example um we have projects like crossplane which kind of extends this infrastructure as code into kubernetes treating cloud uh APIs as CRDs. We have project uh called as I said I'm maintaining this metal cube project which is trying to bring the same declarative model to bare metal provisioning. Then we have um some uh when it comes

to policy we have OPA and CUno which as we know handles the policy enforcement layer. Um and the idea here is to have some sort of a control plane composibility but also keeping the drift uh uh in mind and of course as well as the policy traceability. The tag infrastructure is trying to set some best practices for integrating this infrastructure as as code and policy pipelines. Um

and again the goal is to have zero drift infrastructure which is auditable. We have uh self-healing and of course observable from comet to runtime. And then lastly, edge and sovereignity which is also becoming critical as cloud native systems move beyond these centralized data centers. So at the edge we are extending the cloud native principles into highly distributed resource constraint environments. That means dealing with interminent connectivity, heterogeneous

hardware and workloads that are sensitive to location. And as we saw in the keynotes as well today, the sovereignity side uh our focus is to shift uh to control and compliance. Organizations need to ensure that their systems operate within specific legal and regulatory as well as organizational boundaries. So the idea again here is to have or or have the uh the cloudnative design push towards greater distribution,

resilience and and control. So what are the uh deliverables that we are trying to produce uh according to the charter. So first and foremost we are trying to produce the frameworks and white papers. Dylan is going to talk about some of the ongoing uh initiatives that we have in tag infrastructure space. But again the idea is to have some uh produce these frameworks or white papers or

best practice assessments used by CNCF projects during these sandbox to graduation reviews as well as have reference implementations and um testing guidelines and the success is of course observable adoption across the ecosystem so that and of and as well as projects are also aligned when it comes to APIs or practices with stacks guidance. uh we also want to see new contributors participating in initiatives and there is

some sort of a measurable cross project consistency. Um lastly in the charter we also define how we are coordinating across the CNCF ecosystem. So as you saw in the initial pictures where the the tax was introduced by Dylan. So for example uh with tax security we are uh aiming to focus on runtime hardening and supply chain policy. When it comes to track workloads foundation we try to

uh work with runtime and scheduleuler interfaces and with track operation resilience we are aiming to work on recovery and continuity. Uh one of the most important activities that we also do as part of the tag work is to um engage ourselves in the sub project reviews. So we are also working with the TOC sub projects. Uh and lastly the goal here is to ensure that every tag

and project builds towards a composable interoperable infrastructure stack rather than isolated silos. So I hand over to Dylan for the >> Thanks Kashie. So one of our initiatives that we have on going ongoing right now that kind of carried over from some of the previous tags uh into now is the uh storage landscape v3. So there was a white paper a couple white papers before obviously a

v1 and v2 um but in this initiative we wanted to complete this version and discuss the definitions of the attributes of a storage system. um those being availability, scalability, performance, durability and consistency. Uh in this white paper, we'll describe different storage layers and how they affect the storage system attributes. For example, storage topology can be centralized, distributed, sharded or hypercon converged. We describe how that can affect

a storage system. We also describe the data access interface such as volumes and application APIs. Uh and we also discuss different types of storage systems block file system object stores key value stores databases streaming and messaging. We'll discuss how the container orchestration systems such as Kubernetes interact with storage systems using interfaces such as the container storage interface and the container object storage interface. Another initiative we have

is data storage in cloudnative AI. So this in this initiative this is another white paper that describes characteristics of AI and ML workloads and what that means for data storage. So this is a completely dedicated and separate talk from the V3 landscape uh specializing around um some of the new and emerging needs from AI and ML workloads. We will describe the patterns and trends uh including data

warehouses, data lakes, data cache and locality vector databases and so on. We want to review how storage is being used in the AI life cycle right now. So with training and inference, it's such a new space that this is something that we're going to be uh looking for contributors across the board and getting use cases on. uh we'll be evaluating what are the storage requirements and uses

patterns in each phase. Um when you're doing training you're using large amounts of data and sometimes even high performant storage is required. Um you also need a storage system as I mentioned with high throughput and high capacity. Um for inference you use smaller amounts of data. you need a storage system with low latency and high IOPS um such as flash, MVME or even memory on the GPU.

And lastly, we have an initiative around infrastructure life cycle. Uh one of the one of the things we wish to address with this is something that's been emerging even prior to AI and been accelerated with AI. uh is that over the 10 years uh in cloud native uh infrastructure has been super focused on public cloud and this is apparent in some of the projects like cluster API

and some other industries where how are you what are the patterns around provisioning infrastructure and how you manage that life cycle when it comes to some of the emerging spaces spaces such as edge and IoT and even private cloud and bare metal uh as the push for AI continues uh you see some of these typical cloudnative practices that have been matured in public cloud not apply as

well onetoone with these emerging spaces. Uh so guide but being guided by these cloudnative principles around being secure res resilient observable and manageable. We want to re-evaluate what does infrastructure life cycle mean acro across all of these environments and address some of the emerging trends with crossplane being reconciliation based and some of the more industry norms like Terraform which is more event driven. We al also want

to address the differences between declarative imperative and address some of those use cases that people are seeing when it comes to trying to understand which one to pick and which one best fits your needs. So, I covered three three of the initiatives that are ongoing in the tag right now. If you have any more uh ideas, we're happy to have discussions about any potential new initiatives or

sub projects uh that you wish fits within this domain. Um but some other ways you can contribute to the tag uh is um by contributing to some of the existing CNCF projects. Um you know, we need contributors for initiatives. Even if you think that you're hey I'm somewhat new in the space um uh a fresh perspective is always great uh when discussing these top high level abstract

topics and white papers um we need contributors for the project reviews um the even though the tag leadership the tech leads and the uh chairs uh are very much involved in the project it's always good to have contributors with a fresh set of eyes helping us review projects uh to make sure they meet the bar for the rest of the ecosystem and catch things that maybe we

uh uh already know what to look for. Um refining I mentioned this a little bit earlier about the charter 2 re consistently refining our scope. What should we be looking at that isn't in our scope that maybe is up and coming or things that maybe need a drop off. Um are there any gaps helping us to identify industry trends? Uh what's up and coming? uh and just

be at the bare minimum being community advocates, posting on social media, going to other conferences and talks and uh uh and uh spreading uh the um the mission that we're trying to achieve in cloud native. Uh now this is an example ladder here for folks uh at at any level whether you're an expert in the in your field or you're just starting out. Uh this is uh

a great way to kind of showcase or visualize how you can start contributing to the tag today. Uh as I mentioned before coming in attending the meetings sometimes yeah I know we all have lots of meetings. I know I do and sometimes I don't attend the meetings on a regular basis but that's why we have uh multiple co-chairs and tech leads. Uh we get busy personal life

happens. Um but another great way is to reach out to the projects. You know, let us the tag know how are the projects doing, what do they need, what are we where do we need to step in and vice versa. Contribute to those projects and help them uh uh understand where their gaps are. Um uh once you get more comfortable, uh we've had contributors help lead initiatives

and become tech leads. Uh you can even come come help staff the booth in the project pavilion as well. Um and uh maybe eventually become a tag lead yourself. Now here's some more QR codes with some community links. Um the first one is to the Slack channel and the CNCF Slack. Um and the other second one is a link to the uh meeting calendar on the Linux

Foundation LFX uh site. So, I'll leave that up for a few seconds here until I don't see any phones up. All right, everyone good? Cool. And finally, um, one last QR code. Um, that completes our talk today. If you have any feedback about the talk or about the tag in general, please take the time to fill out, uh, the feedback form. Uh and thank you for coming.

We have roughly 3 minutes. If you have any question, please stand to the mic. Otherwise, we are also available in the booth. >> Hi, I have only a teeny tiny question. Uh maybe I missed it at the start because couldn't find the room. Um but um is this also where the reference architecture um for the CNCF essentially is also being uh integrated because part of it is

of course uh white paper standard stuff like that and the reference architecture are I suppose more in the corner of case studies and how people are using it but it seems like they would essentially complement each other. Is that something that's happening within uh the same scope or is that completely separate? No, it it it definitely either it's I don't think it's something that in terms of

reference architectures um I know that some tags have put out reference architectures before. So when the the tag app delivery that existed before that the platform working group came out of one of the artifacts of the platform working group was the uh platform maturity model which had a reference architecture for this is how you build platforms. This is how you use some of the cloudnative projects to

build those. So it definitely can be one of the artifacts of the tag or if it starts in an end user or in a community group uh the tag can lend some of the uh contributors or tag leads or chairs to provide technical experience to help develop that architecture. >> I think there is some sort of an enduser uh sub project or something going on and they

also have these sort of reference architectures there. >> Yeah. And the end user one is also the one that's on the like the CNCF reference architecture page that is more I say it's more publicly marketed I suppose that's maybe the best uh so that would be yeah that would be separate yeah it makes sense it's more end user oriented but yeah I was trying to find what

the relationship between all of this uh >> is but yeah that makes a lot of sense thank youly