Sponsored Keynote: Open Source Infrastructure for the AI Native Era - Jonathan Bryce
About this talk
This talk discusses the growth of the PyTorch community and its relationship with the Cloud Native Computing Foundation (CNCF), which oversees Kubernetes and numerous other projects. The speaker highlights the significant shift from training AI models to mainstream inference within organizations, emphasizing the importance of continuous training and updates in AI workloads. They mention how companies are increasingly using inferencing power, expecting a major shift in compute usage from training to inference. The speaker also points out the need for improved models and efficient infrastructure to support this evolution, citing Uber as a notable example of leveraging these technologies. Additionally, they introduce LLMD, a new project within the CNCF focused on enhancing AI model serving, and encourage attendees to learn more about it.
Full transcript
It's awesome to be here and awesome to see the the growing community around PyTorch. Um I I I Mark mentioned, you know, that we we started OpenStack. And so, we actually have been able to see this journey on on a few projects. And it's it's uh I feel like we're just at a really early point still in this community, even though it's growing. I think there's there's
a lot more to happen. That's very exciting to be part of it. So, um how many of you are familiar with the the Cloud Native Computing Foundation, the CNCF? Okay, most. All right, good. Um how many of you know Kubernetes? That's uh yes, everybody Okay, this is uh this has been my experience this week cuz a lot of people don't really know the CNCF, but they know
Kubernetes. And we're the the foundation uh behind Kubernetes and and uh and about 230 other projects in the space. Um so, really incredible global community of people. And what uh what we focus on in in this community is how do we take technologies and scale them operationally? How do we make it so that every company in the world can run different kinds of technologies? And right now,
what uh what we've been seeing in our community is obviously a huge uptick in platforms that are um implementing different types of AI workloads. Um Mark mentioned yesterday kind of what we we think about as as three pillars in open source AI with training, inference, and and agents. Um and this is a thing where uh you know, each of these these types of systems are uh they're
they're all scaling more than they ever have before. They're also the lines are blurring uh with the way that we are seeing continuous training techniques and and uh and model updates. And this is creating a really interesting inflection point for um the broader technology landscape that we see very much in in the CNCF. And I think the biggest shift that we are seeing right now is a
real uh move from just kind of massive massive training environments to mainstream inference um that's happening in in almost every company that I I talk to. Uh and I think we've hit these these um moments over the last couple of years crossing the chasm, so to speak, where uh you know, this is not just a kind of like a science project anymore, but something that's implemented in
everybody's business. Um Octave Klaba, who's the CEO of OVH, uh he said, "We're all waiting for tokens." And uh and I think, you know, that's true. Like, we're all we're all waiting for how can we use intelligence faster? And there's really two ways to do that. You we need faster models and answers, and we need more inference. some data that came out recently, specifically around inference, is
that uh this year the kind of makeup of AI compute usage is flipping from being 2/3 training and 1/3 inference to being 2/3 used for inference and 1/3 used for training. And by the end of the decade, there's going to be 90 plus gigawatts of compute power online for training. Or or sorry, for inference. And that's more than all other workloads combined. So, we we are seeing
this just massive investment, and it's driving a lot of adoption. And I think the other thing that's really interesting is we're seeing a different way that models are being developed and used. And we heard from Ray earlier, who's one of the pioneers in in these techniques. Um Uber is a great use case for this. They're a big user of Ray and and Kubernetes. Um but they are
seeing how by continuously improving their models, they're getting better answers faster. They're making better utilization of their hardware. And they are able to achieve um a much better result for, you know, for their business and their developers and their teams. So, uh what I really just kind of want to leave everyone with today is that over in the CNCF, we see inference, and especially this kind of
new form of it, as uh as a huge workload in the cloud native landscape. And we really want to make sure that we are working across communities so that we can deliver um the best infrastructure, uh make the most usage of of this uh investment, and uh and really think about like how do we build these systems so that we can scale them to every organization. Um
we just added a project in the CNCF called LLMD, which is uh a a great platform built around vLLM. There's a talk today at 11:05 in the central room around LLMD. So, encourage you to go check that out and uh and um learn a little more about that. And um you know, ultimately, I think as this community continues to grow, we'll see um you know, bigger events,
more projects, all of those those awesome metrics. But what's really going to matter is, you know, can we build the systems that deliver on the promises? Everybody out there, you know, we all want agents, but what we need is we need great models, we need inference uh platforms to serve them. And uh and that's how we deliver intelligence for everyone. So, um thank you. And uh happy
to be here at PyTorchCon.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17