KubeCon + CloudNativeCon Europe

Keynote: Welcome + Opening Remarks - Jonathan Bryce & Chris Aniszczyk (ASL)

44:35 · 23 Mar 2026 – 26 Mar 2026 · YouTube

About this talk

This talk marks the opening of KubeCon Amsterdam 2026, emphasizing the theme of advancing Cloud Native technologies. The speaker highlights the growth of the Cloud Native community, mentioning over 13,500 attendees from more than 100 countries, and the expansion of the Cloud Native Computing Foundation (CNCF) with over 230 projects. The discussion includes insights into the significance of AI in today’s infrastructure, with a report indicating nearly 20 million cloud-native developers globally. Recent project advancements are noted, including the graduation of Kyverno and Dragonfly, as well as the introduction of end-user reference architectures from companies like Swisscom and CERN. The speaker also welcomes NVIDIA as a new platinum member supporting CNCF and discusses various AI infrastructures, emphasizing open source as the foundation for future AI advancements.

Full transcript

Good morning and welcome to KubeCon Amsterdam 2026. We are We we are glad to have you here. Um the theme for this KubeCon is all about let's keep Cloud Native moving. And so to kind of kick things off uh let's move a little bit. So if this is your first KubeCon Europe, stand. Please for So it's that's a lot. That's about 40% of you or or or

so. Uh thank you for for kind of joining on this on this Cloud Native open source uh uh journey. So today we have uh a little bit over 13,500 uh attendees making it the biggest KubeCon Cloud Native uh con event that we have done which is awesome. Thank you so much. Uh represents a little bit of 10% growth over last year. We have over 100 countries represented

here uh for you know for attendees over 3,000 unique organizations coming to learn about Cloud Native. And we have about 900 sessions for all y'all to kind of learn and nerd nerd nerd out about uh this week. And we have a one glider. And we'll and we'll get to that a little bit later. as you could tell the community is is still growing. Uh it's incredible. We

have over 230 projects in CNCF, 300,000 uh contributors worldwide that are contributing to our projects. And you know we continue to grow. You know a lot of people may not realize but our modern digital infrastructure for almost everything you use from taking a train, calling in a calling an Uber um you know going through airport security. A lot of these technologies power or powered by uh CNCF

projects. And we're currently in this uh you know I think what what Jonathan and I refer to as like we're in the AI and agentic you know era. And with this there's more development, more services, more infrastructure to manage. And of course our community is growing along with that. So we have a report that we published at the last KubeCon and we did a refresh in you

know, last 6 months and now we if you look at the total amount of cloud native developers out there, we're almost at 20 million. So, it's amazing to kind of see what our community is coming together to grow and help truly not only build the modern digital internet, but also power the AI infrastructure revolution that we're kind of going under. So, if you're interested in this report

and learning a little bit more about details of the cloud native development community, go check this out. There's a lovely QR code here for you to poke around and learn The other thing to mention is you know, we are in Europe. There's a lot of folks that truly care about digital sovereignty here. We had our open sovereign cloud day yesterday at KubeCon and sometimes there could be

misconceptions that you know, CNCF you know, there's maybe one country or one one company or one organization that does the majority work, but that's actually not true. We are truly a global and diverse organization and we're putting together some data. John and I were poking around and we're actually a little bit surprised. You know, Europe makes up the largest you know, dominant contributor across all of our

CNCF projects. So, super impressive for all y'all. So, give yourselves a hands. And and of course, you have you know, the US and India, China, Japan and many other parts of the world that come together and do this. So, if you're interested in poking around with this data, we have an excellent developer dashboard tool called LFX insights. You can go scan that QR code or go to

insights.lfx.dev to go play around. we continue to grow. We've expanded in terms of community and cloud native developers. We've also continued to expand the cloud native project landscape. We've graduated some recent projects in our community. We have Kyverno which graduated which is a great project focused on security and policy in the cloud native landscape. You have Dragonfly which is really all about distributing large binaries, containers, and

even AI models at scale which is awesome. We have fluid and Tekton, which entered the incubation level, which are focused on continuous delivery and accelerating AI usage in a Kubernetes kind of context. And we continue to add innovative sandbox projects that focus on making Kubernetes an amazing place to run gaming servers, and you're going to hear about that in some of the keynotes later today from, you

know, making inference run amazingly in a Kubernetes. So, there's kind of continuous innovation that is happening in CNCF, and you could always go to the fun landscape that we have at landscape.cncf.io to kind of check out all the innovation that is happening in We also have some new members that continue to join us, which is always amazing. So, I'm happy to announce that we have two new

gold members that have joined us recently since the last KubeCon. We have F5 and Viettel. So, let's congratulate them for joining and supporting our you know, ecosystem. Great, great organizations. Um and of course, we have a lot of other new members, end users, nonprofits that can continue to join and support CNCF. One great thing about our community is we have a lot of end user companies that

are involved from, you know, telecommunications companies to automotive to train infrastructure. And one thing that we love to do is to we love to work with our end user community and develop essentially what we call end user reference architectures. So, how do you actually deploy all these projects that do something cool, right? And so, we have some new ones that we are announcing today where we have

a reference architecture from Swisscom of how to deploy a sovereign cloud that they're using. We have Zeiss, which kind of built a really cool order fulfillment system. And then, CERN, of course, who has done a really cool AI, you know, scientific-focused cloud. So, you can go check these out on architecture.cncf.io. I encourage everyone out there, if you work at a cool company and you're doing some cool

things with CNCF projects, go submit these end user reference architectures. KubeCon and the CNCF community is all about learning from each other and sharing lessons. So, check them out. We'd love to see more more of these. you know, as as I talked a little bit about it it's it's truly hard if you work in software engineering and and kind of work in technology these days, it's almost

impossible to escape uh AI, right? It's it's it's all around us. Um you know, I use both, you know, AI tools, Claude, like to help me with coding assistant and also as a therapist. It's great great great great tools. Um but you know, this is what the new world is going to be about. We're going to integrate these things everywhere and you know, I'm really excited to

in CNCF that one of the companies that kind of has been at the forefront uh of of this AI revolution and one of the largest uh and most valuable companies in the world uh has decided to join and support CNCF uh at the platinum level and I'm very excited to welcome uh Nvidia as our latest and newest platinum uh member. So, let's give them a grand hands

and personally uh really excited to go uh welcome an old colleague uh of our of ours on stage to talk a little bit about why Nvidia is joining and supporting CNCF and open source. So, let's go uh hand it off to Aaron Boyd to come on stage and talk a little bit about Nvidia and what they're doing. Good morning. How's everyone doing? Get some coffee? Awesome. Yeah,

so the future of AI is community-driven and open. Now, let's look um a little bit in the past to think about where we're going and think about the technologies that have reshaped computing. The operating system, the internet, virtualization, the cloud. But Kubernetes started smaller. It was a scheduler. It was a way to run containers across environments reliably. A tool for teams dealing with scale. But then something

unexpected happened. Developers didn't just adopt Kubernetes, they standardized on it. Operators and service providers didn't just run it, they built platforms on top of it. And organi- organizations didn't experiment with it. They bet their business on it. And somewhere along the way Kubernetes stopped being just infrastructure and became the de facto programmable control plane for modern distributed infrastructure. And here we are today, over 13,000 individuals who

are part of an ever-evolving ecosystem of open source engineers who are running mission-critical systems, globally scaled services, and increasingly, as Chris mentioned, AI workloads. And that's why it keeps expanding, from containers to databases, from stateless to stateful, from apps to platforms. These AI workloads are driving innovations in infrastructure that Kubernetes wants to standardize and adopt and simplify operations, but also enable the developer like never before. Nvidia

has been part of open source at the lowest levels for many years, but in the past 4 years, we started to accelerate our development to open And more recently, contributing directly to the cloud-native ecosystem. Schedulers like Kai is officially now in the CNCF sandbox. Woo! Yeah. Last week, we open-sourced AI cluster runtime, or we affectionately call it ACRE, because as engineers, we simplify everything, which delivers runtimes

tuned for AI. We are able to verify the configurations are conformant by running it against the Kubernetes AI test conformance test suite for inference using distributed infrastructure workloads like Dynamo, which just was released to 1.0. Anyone who is using an NVIDIA GPU today is definitely using one or more of our tools within their stack. But I am most excited to announce today that we are donating the

NVIDIA GPU driver directly to Kubernetes SIG Node. The driver provides a reference implementation for the vendor neutral Kubernetes DRA API, which helps standardize AI infrastructure. But you know what developers need most? Access to compute. And as part of our commitment to developer enablement and the cloud native ecosystem, we are pledging $4 million over the next 3 years to ensure that all projects in the CNCF that need

NVIDIA GPUs will get access to them. The future of AI will be built in the Just like Kubernetes is. Why? Because the hardest problems ahead are not just model problems, they're infrastructure problems, scaling problems, interoperability problems, trust and transparency problems, and no single company can solve those alone. Open source is how we share innovation, standardize platforms, build ecosystems instead of silos, and it's what made Kubernetes the

foundation of modern infrastructure, and it what will be makes AI the foundation of the next generation of compute. Thank you and have a great KubeCon. Great job. Thank you, Aaron, and I just, you know, I think it's so awesome and exciting to have NVIDIA joining as a platinum member and then also supporting the community with that $4 million grant to help us test on on the latest

hardware that's going to be put to good use, I'm sure. Um I want to talk a little more about AI 4 months or so ago, we had KubeCon North America in Atlanta. Um and there I talked about uh where I thought we could really have an impact in the world of AI with our cloud-native community and our projects and technologies. I talked about three pillars, training, where

we take data, we encapsulate it, and uh kind of put this intelligence into a model. Inference, where we take that model, we serve it, so we can make predictions and answer questions. And agents, which are really where we see that intelligence become accessible uh to humans, to other agents, to software, to systems. And this is something that's um you know, I think is is amazing to see

how much progress there's been in just those 4 months. At that event, uh I I talked about inference all week. Uh I, you know, had so many great conversations. A lot of those conversations were talking about, "Well, I've been trying to put together an inference system." Some of those conversations were, "What is inference? How does it work? Why is it important?" The conversations I'm having now, though,

are really changing in a significant way, and a lot of times the conversations now are, "How do we scale inference?" And this is not just something within our community. If we look at the global macro environment around IT and AI, we're seeing this in all of the areas that we can measure. Um if you look at how AI compute was distributed in 2023, 2/3 of it went

to training, and 1/3 of the compute went to inference. And by the end of this year, that's actually going to flip, where 2/3 of AI compute is going to be dedicated to inference And by the end of this decade, which is just a few years from now, uh the amount of uh compute capacity that's going to be dedicated to inferences is over 90 gigawatts. 93.3 gigawatts. And

this projection is actually that that will be greater than all of the other compute loads combined. So, when I say, you know, inference is going to be the biggest workload in human history, it's not an exaggeration and it's not something that's 10 or 20 years away. This is happening right now. And we also see that this is driving new markets and new opportunities. I think that's what's

awesome for our community is we have a huge growth potential ahead of us for our projects, for our companies, for our developers. Uh and this is something that you we can act now to help take advantage of this. And so, what is driving this? What's driving this insane demand for Chris mentioned Cloud Code earlier. How many of you have used Cloud Code or Code X or Open

Code or Goose? You know, one of those coding agents. Exactly. Uh what what about Open Claw? Is anyone willing to admit they've got their claw running? Yeah. Okay, we've got a few brave souls out there. Well, you know, when you look at agents, these are super users of inference A lot of times people think about AI as kind of the the chatbot experience of chatting with uh

with ChatGPT or something like that. But when we move to the agentic world, the usage skyrockets many, many multiples. The good news for us is these are the expertise that we have in this community already. How do we take distributed systems, deploy them, observe, scale, secure? These are the right skills that we have and this is exactly what the AI world needs right now. And I think

that the other good news that we can look at, we have a great start in this. Uh we have a huge footprint of Kubernetes in most of the organizations in the world and many of those organizations are already running inference workloads on top of And as Aaron alluded to, you know, this isn't because Kubernetes was built for AI and inference. It was built for distributed systems and

this is becoming the most widely used distributed system in history. I think as we look forward, what we need to be thinking about is how does cloud native support the new and right here AI native era? How do we take the primitives that we've already created and built and expand on them to support AI workloads? And to do that, we're going to need updates to existing technologies.

Um we've heard about DRA today. Uh the the inference gateway that's in Kubernetes I think are great examples of where our projects are progressing. And we'll also need new projects and new technologies. When I was in Atlanta, I mentioned an example of of a project like that called LLMD, which is a distributed inference uh system built around Kubernetes. And what I'm really excited to talk about and

announce today is that LLMD has just entered the CNCF sandbox. So, please help me welcome uh the team who's working on LLMD to tell us a little more about this. Check. Hey Rob. So, we've heard Chris talk about inference or sorry, we've heard uh Chris, we heard Jonathan, we heard Aaron. Um and we're just getting started, right? So, can you help us understand more about why inference

is so important right now? Thanks, Karina. Let me start from the beginning. When ChatGPT launched in late 2022, it was honestly a bit of a scary moment for the open model ecosystem. Proprietary API-based models were simply an order of magnitude better. However, since that day, we have seen an absolute explosion of capability in the open ecosystem. Starting in 2023 with the first models like Bloom to 2024's

Llama 3 moment, the first usable open-source model to 2025 with the introduction of large reasoning MOE models like DeepSeek and NVIDIA's NeMo Megatron, the open model ecosystem is simply on fire. However, all of this progress creates a huge challenge. How do we scale inference against simultaneous 100x increases in model size, context length, and request-level token intensity? As a lead developer of vLLM, we knew that we needed

to invest in the next frontier of optimization to deal with these challenging agentic workloads, distributed inference. And so, while vLLM optimizes a single node, squeezing as many tokens as possible out of every host, LLMD optimizes a whole cluster of vLLMs, implementing distributed performance optimizations like LLM-aware load balancing, KV cache management across the storage and memory hierarchy, prefill and decode disaggregation, and multi-node expert parallel deployments. Thanks, Rob.

Abdullah, now okay. Load balancing, routing, how do they need to evolve to really support all of So, traditional load balancers were built for the stateless web. Routing traffic based on simple metrics such as number of open connections or round-robin. But, LLM serving breaks this model because it's inherently stateful. If you route a prompt to a random GPU, you significantly reduce your chances of KV cache reuse. So,

combined with the widely unpredictable compute costs of auto-regressive decode and variable sequence length, traditional routing practically significantly guarantees cache thrashing and stranded GPU capacity, leading to higher request serving latencies. To fix these inefficiencies, Inference Gateway, which is an LLM D component, it uses LLM load and prefix aware load balancing. Inference Gateway inspects incoming prompts and sends them to the specific GPU that already holds a significant part

of the context in its KV cache. So, to achieve this allows us to achieve much better load distribution, much more uniform KV um utilization, but it also significantly reduces uh um prompt compute and and and slashes time to first token. But, not only that, it also frees up critical high-bandwidth memory on the accelerator, which allows us to have much bigger batches during decode, resulting in much higher

throughput. Thank you. This really sounds challenging. And uh thank you, Matthaeus, for being here with us. Um can you give us a concrete real-world example from Mistral AI on how challenging this is? Yeah, so one concrete example we wanted to share is disaggregated serving. Um the core idea is simple. You split the prompt processing phase, also called prefill, which is compute bound, from from the token generation

phase, also called decode, which is memory bound. Each phase run on its own little worker set. This approach is becoming essential, uh especially to serve large mixture of expert model like Mistral large 3 or Mistral small 4 that actually got released last week. Um so, by adapting parallelism strategies, you directly improve your MOE throughput. And on large deployments, even even on four dense models, it stabilizes the

decoding speed and reduce your tail anti-token latency. So, your quality of service improves using fewer GPUs. Yes, the deployment is more complex, but the improvements in performance are hard to ignore. At Mistral AI, we think that open collaboration on these issues is essential to building future-proof infrastructure and setting new standards. For example, we identified the need for a disaggregated site operator that will be released within the

leader worker set project. Essentially, it synchronizes the prefill and decode phase by creating a tightly coordinating rolling update. With this, frameworks like LLND can ensure only compatible versions communicate with each other, making such deployments safer to deploy. Amazing. Thank you, Matthieu. Now, Carlos, yeah, you can clap. You can Right? And Carlos, let's let's bring it home. Can you help us understand how this is all coming together

and helping solve these challenges in a production environment? >> Absolutely, Carina. Thank you. So, these are really cross-layer challenges. They cannot be solved independently. That was the motivation for the creation of LLND. So, last year, we brought together a coalition of industry leaders with one common goal, to make inference a first-class citizen in Kubernetes. The design principles were simple. Make it composable by building open standards like

gateway API inference extension. Deliver proven paths from experimentation to production for state-of-the-art inference. And make it flexible. Run any model on any hardware and any cloud. We've come a very long way. We've proven ML we serve in this aggregated a production reality with significant performance gains. That's leading ML MD to be cited and adopted in leading industry forums. And that's why we're taking the very next big

step. We're very pleased to announce that ML MD is joining the CNCF Foundation. We truly believe this will be a catalyst to grow the ecosystem and with the help of you, the community, tackle the next big frontiers. Be it workload aware, KV cache, multi-tier caching, scaling reinforcement learning, and doing autonomous configuration. So, come build with us the future of Thanks so much. All right. so I I

think that that's going to be a very, very important project for enabling all of the incredible workloads that uh that we see out there. Um and I'm I'm so happy to uh to have it in in uh the CNCF sandbox and see where we can all take it together. There's another trend that I think is is starting to emerge that's also very interesting to me and that's

around specialized intelligence. Uh I think the last few years we've seen a lot of uh dominance from kind of the uh the the foundation models, these frontier models. And a lot of times, you know, that's still what people think of as AI is that kind of chatbot experience. But this next phase of AI, which has already started, is going to be a little bit different. I think

we're going to tens of thousands, hundreds of thousands, even millions of models which get embedded into all kinds of environments. and rather than just the pre-training on a set of public data like a lot of the this is where we're really going to see the value of private data unlocked by taking those foundation models and post-training them, fine-tuning them, by doing reinforcement learning and by doing continuous

training. And so I'm really excited about specialized intelligence, specialized models as something that's going to again drive a lot of investment, a lot of infrastructure usage, and a lot of advancement in what we can continue to get out of out of AI. And you know I think there are are advantages to this too that that are really going to matter for people because these models they can

be more cost-effective to run, they can be faster, they can give better answers and all of that means that you have options for where you deploy them, how you use them, what the security posture is. So I think this is going to be another big trend that we're going to see and it's going to impact our projects and our work and and what we need to to

think about building and and how we run this. I'm not the only one who who's thinking in this direction. Clem from hugging face posted recently on uh on X I guess. I was going to say Twitter. Um but uh he he talks about this how you know we we can we need to move past this misconception that training is only for those who have billions in resources

and and so many GPUs that that that they can you know do these If you go to hugging face you can look through and they have so many interesting models that you can use directly or that you can start from and then you know fine-tune it make it work for your use case and I think that this is what we're going to see. Now in order to

do that you know we're going to need to improve the pipelines for our training and our inference systems to to to be able to do this continuously on a regular basis. One of the companies who's doing that today and I think is an absolutely incredible example of this in the real world in action is Uber. And they're here to tell us a little bit about their world,

their use case, how they approach this. So, please help me welcome Melda Salhab. All right, thank you. I'm very happy to be here. I'm Melda. I'm part of the Uber AI team and I'm here to talk to you about my Michelangelo AI, Uber's ML platform. Just going to speak a little bit about how we use AI ML at Uber and also fundamentally how Kubernetes is a critical

piece of that puzzle. So, first, how does Uber use AI? So, AI has, you know, been a thing for the past kind of buzzword for the past few years, but ML or AI has kind of been a fundamental part of the Uber experience from our very inception. So, each time you interact with a product, there are many, many ML models running in the background. So, that obviously

powers things like our marketplace, so dispatch, pricing, matching, so on. And then personalization. So, if you open your Uber Eats feed, you might get a different view than someone else. Then also kind of very fundamentally other parts of the platform you might not be aware of like our risk and safety. For example, we want to make sure that the person picking you up is indeed the person

who signed up on our platform. So, AI ML kind of powers a lot of that. And then with GenAI recently, we've now integrated quite a lot into our platform to just take things to the next level. So, what's what's the challenge? What make Why why are we here today? So, fundamentally, the big challenge at Uber for ML in one word is scale. So, we are live in

over 70 countries, over 10,000 cities, and that translates to actually well over 33 million trips per day. Then, if we take that down to the ML world, that means we as a platform have to support well over 30 million peak predictions per second across a thousand serving nodes including CPUs and GPUs. So, how do we do that? Michelangelo is the Uber ML Ops platform. We've been building

it since 2016. So, as of now it's about 10 years. I'd say it's our 10-year anniversary. In the early days, it was all about linear tree-based machine learning models. We had a simple UI that kind of abstracted things away from ML developers. We focused on our feature store, workflow orchestration, and quality. From 2020 onwards, deep learning became, you know, the the new technology. And then we had

to both support machine learning, support the more complex use cases, so making sure our developers could use a code-first way of model development, but also just maturing our platform. So, that's where we introduced model performance, feature monitoring to make sure that we can manage this at scale for Uber, and also tearing and just other critical parts of this of this platform. And then from 2023, GenAI and

obviously more recently agentic AI became a thing. This is a new flavor of AI for sure, but fundamentally the challenges are kind of the same. Building a proof of concept is relatively easy compared to taking a GenAI application, agentic AI application, and doing it at scale. So, the same team has been kind of working on how do we do that? Different things like model gateway, LLM serving,

and then now we've done a lot on the agentic side, which I'd be happy to talk about separately. So, what this means is Michelangelo has been able to support 100% of our mission-critical ML at Uber. Translates to 20,000 models trained per month. 5.3k of those are in production, so live hitting the different trips every day. The 30 million peak predictions that I mentioned, and we operate with

serving reliability of four nines. So, how do we do that? This is a very, very short and simple representation of Michelangelo. So, at the fundamental, we have the data plane, and we also leverage a lot of open source technologies there. So, from PyTorch, Spark, Ray, TensorFlow, so on. And this is across both CPUs and GPUs. And then, fundamentally, our control plane is where Kubernetes is the is

the critical piece. So, we have a Kubernetes-based API, and that just manages everything. Make sure that we're doing it we're achieving our reliability goals. It just manages the workload across different compute clusters. Be happy to speak about that further, so please come to our booth, and we we can share our reference architecture of Kubernetes. And then, most importantly, what this means is that all that infrastructure is

just abstracted away from the developer. So, as an MLE or a data scientist using Michelangelo, all you have to focus on is the machine learning. So, with that, I'm going to wrap up. Thank you for having me. We'd be more than happy to chat more about this at our booth. And yeah, thank you to the team. Thank you, Melda. Uh just to reiterate a couple of things

that I love about that story. 30 million predictions per second, 20,000 training runs, 5,000 models. Uh I think that that is one of the best examples that I've seen of truly an AI-native company. And I love how how much is built around open source, not just CNCF projects, but PyTorch and vLLM, which are in the the PyTorch Foundation, and all kinds of other tools that they bring

together to enable that. So, I mean, one of the things that CNCF is is famous for, and and some people may even take for granted, is, you know, Kubernetes has evolved over the last decade, and truly is supported on every major public, private cloud out there, you know, in every almost geography. What we've done essentially for Kubernetes, for traditional cloud native workloads, we are doing the same

thing for AI workloads. And last KubeCon, we announced the launch of our certified Kubernetes AI conformance program, which is going to go replicate some of those features around ensuring Kubernetes is consistent across different platforms, clouds, and so on, but for AI workloads. And so, we had an initial batch of companies representing some of the largest clouds, and even us, you know, smaller neo clouds out there that

were part of this program, and we continue to make progress. And it wouldn't be a KubeCon without a fun live demo, so I'm excited to bring on Janet Kelner from Google to do a little demo of kind of what's next for AI conformance. So, I'm going to show you what's new in With the importance of inference, and how to scale inference, the community decided to add a

few new requirements in AI conformance around inference. And let's see how it looks. let's bring up the demo. So, let Let warn you, this will be a live demo. Yay. So, I'm going to use a script to help me do all the typing cuz I'm not very good at speaking typing, but let's see. So, this demo is running on a AI conformance cluster. The first requirement I

want to show you is that it must support gateway API for advanced traffic management. That's the load balancing bit you just saw. So, first I want to verify there are gateway classes in the cluster. So, gateway classes show me all the um available gateway classes in the provided by the infrastructure and the controller implements the gateway. And then I want to look at the namespace to see

if there's any gateway available So, there's one inference gateway there using the L7 external um gateway class. Moving on to the next requirement. AI conformance platform should support gateway API inference extension for advanced inference routing. This is for um inference aware load balancing that you also just heard. We first verify the same gateway that I just showed you. We want to see there are routes attached to

it. So, I can see that there are two routes. And those are actually HTTP routes because we're using inference extension. It's um routing traffic to inference pool instead of a service. That's how it's inference aware. So, as you can see, I have those both routes using the same inference pool. So, what is an inference pool? An inference pool defines the set of model serving pods behind the

gateway. It also has the endpoint picker that will route the inference request to the optimal model serving pods. So, something like KV cache aware routing that's handled by it. And the last requirement I want to show I want to show you is that the AI disaggregated inference. This is about scaling inference. It breaks the prefill and decode uh phase into separate scalable components. And then, you can

scale scale them differently and scale them on different a dedicated hardware pool. So, let's look at the model serving pods in the First, we have the prefill pods. They are there for handle handling the prompt processing, which is very compute heavy. And we have one prefill pod. Next, we have the decode pods for handling the token generation, which is very uh memory heavy. So, we can scale

them differently depends on our need and put them on different nodes. So, we I have two uh decode pods. And this example uses uh LMD, but the platform is free to choose any other disaggregated inference solution they want to support. And the community is also exploring the idea of having a common API for inference. So now we have all the layers in place that see it end

to end. So I'm going to send a live inference request. It's going to be sent to the gateway, to the inference extension for the inference routing, and it will be routed to the pre-filled pod, and then be routed to one of the decode pods, and finally it will return me a response. So I'm going to ask a pretty big model a question. Tell me how Kubernetes help

run AI workloads at scale. Okay, it tells me why and how. And that gives me confidence that Kubernetes is the best platform for running AI workloads. Let's go back to the slides. So um I'm very happy to be here to share with you the new advancement in AI conformance. If you're interested, please scan the QR code or take a picture to get certified and design and contribute

to the Kubernetes AI conformance with us. And thank you so much and enjoy KubeCon. So just as we're kind of wrapping up this this first segment here, I uh I want to talk a little bit about how important it is that open source really does win for the AI era. If you go back just a couple of years most of the AI work that was being done

was being done not in the open. There were open tools underneath like PyTorch and other things. But what we've seen in just the last couple of years is a true acceleration of open source at all layers of the AI stack and I think that is so critical and this is why it's great that you're all here this week working on this. I truly believe that AI infrastructure

has to stay open because this is going to be the intelligence layer for all of us and for the world. And if you go through this week, you'll see that we have a lot of experts here already that we can all learn from and that's how community works. None of this is going to happen without a great community. I also encourage you to look at other open

source communities in the AI space because AI is huge. And so within the Linux Foundation, we have the PyTorch Foundation doing great work at the low-level layers of AI. We have the Agentry AI Foundation building MCP protocol and other other kind of Agentry type technologies. And so join all of these communities so we can make sure that open source wins. We got a lot of content this

week. You know, look on your schedule and see. These are just a few highlights where we've got GPU ops and Agentry ops and more more and more about Kubernetes as well. But I think that this is so important. I'm so happy that we have our biggest KubeCon ever here while we are poised to go take advantage of this opportunity and make sure that that AI does stay

open. Thank you, Jonathan. I definitely think that's super important. You know, our community has grown and there are so many different ways to contribute both from a code level to even a non-code level. We have a great amount of set of projects. You could contribute upstream. We have a technical technical advisory groups that TOC who's responsible for a lot of technical decisions. All these meetings are open

and available for you to kind of learn and join from. If you're not a developer, we have a lot of non-code contribution options, too. You could host a meet-up. You could contribute to a variety of different documentation. We have a glossary. There are many, many ways to contribute within our ecosystem. We have a great ambassador program. Many ambassadors are here on on stage. I've seen There There

they are. Find them. They're very nice. They'll help you out. We have our kind of educational ambassador program, KubeSpenauts. We have about 500 or so in Europe. These are folks that are domain experts in a variety of different aspects of Kubernetes. So, go find them. They have these cute little blue jackets that see. We have a variety of mentorship programs available that we do a few times

a year that we actually pay you stipend to contribute. And of course, we have meet-ups all over the world on community.cncf.io. So, there are many, many ways to contribute. A lot of new folks here. Make a friend. Learn something new. That's what we're all about here. we do these events. KubeCon's are fun. It's very little stressful sometimes planning these, but we're excited that four other KubeCon's coming

up this year. We have KubeCon India, you know, for the second time up in Mumbai, June 18th-19th. Oh, I got some fans. We're back in Japan in Yokohama, July 29th-30th. Yeah. China with our friends from Open Infra and PyTorch doing a little combined story with them in Shanghai, China, September 8th and then back to Salt Lake City for KubeCon North America in November 9th-12th. We plan these

things in advance. So, we want you to save the date because I know all y'all love KubeCon. So, we're going to be in Barcelona next year. Back there. It's going to be great. New Orleans, Louisiana for North America 2027. A little closer to home for me, which is going to be great. And then we're coming back to Berlin in 2028. So, we'll hope to see you there.

Save the date and you know, Jonathan and I both started, you know, this this event with the whole theme is let's keep Cloud Native moving. We're in a whole AI gigantic revolution. We're going to evolve and kind of continue and build infrastructure that everyone depends on in the world. So, thank you for being here and have some fun and learn something new. We'll get on with the

Yeah, we're going to So,