PyTorch Conference Europe 2026

Sponsored Keynote: Any [ Agent | Model | Accelerator | Cloud ]... Maryam Tahhan & Nicolò Lucchesi

4:15 · 07 Apr 2026 – 08 Apr 2026 · YouTube

About this talk

This talk focuses on Red Hat's strategy and initiatives in the realm of open-source artificial intelligence. The speaker discusses Red Hat's commitment to open-source development, positioning the company as a leader in projects like vLLM, which serves as a critical tool for optimizing AI inference across various hardware accelerators. Red Hat emphasizes the importance of high-performance AI, particularly in the context of enterprise-grade solutions, and highlights collaborations with CPU vendors like Intel, AMD, and Arm to enhance CPU integration with AI models. Furthermore, the discussion covers the development of the vLLM CPU performance evaluation tool, a project that invites community involvement to standardize benchmarking practices in AI inference.

Full transcript

I'm Nicola, and I'm here today with Miriam to tell you a bit about Red Hat's work and strategy around open-source AI. So, Red Hat is the open-source company. And we really work upstream. So, Red Hat actively participates in and fosters open-source community with an open-source development model. That is, we do most of our software development directly upstream and then integrate upstream projects into our products. Bundling them

with a an ecosystem of services and certifications. And in fact, more than 90% of the Fortune 500 companies already use some of these products. Including, but not limited to Red Hat Enterprise Linux, which is what Red Hat is famously known for. you may actually recognize some of these projects as part of the PyTorch Foundation, such as vLLM, where Red Hat is currently the leading contributor to the

project. And others, such as LLMD, Smart Router, or got LLM, where Red Hat is the founding member. taking a step step back, in the early days of Linux, enterprises had to be committed to a single vendor when choosing an operating system. And similarly, in the AI inference space, we're enabling running your model across different accelerators to a common stack, bridging model and hardware boundaries. And this is

where vLLM sort of emerges as the Linux of GenAI inference, by allowing to broadcast other aware optimizations to past, current, and future generations of large language models. And speaking of inference, Red Hat is actively developing on the latest generation of high-performance data center GPUs, such as the Nvidia Blackwell family to to ensure future-proof performance across new silicon. we achieve this by optimizing vLLM both system and CUDA

level with model and hardware specific kernels and optimizations that get better with every release. So the message is that vLLM really does not trade off in performance. Enterprise grade open source AI is also high performance AI. Thanks, Niccolò. Um thanks, Niccolò. Apologies. So earlier Niccolò showed us the incredible optionality we have with the latest and greatest accelerators. But here in Europe with the new energy efficiency directive

and great constraints, we cannot simply go out and build new facilities and pack them with accelerators. So we must take advantage of the existing hardware that's already powering your infrastructure today, the CPU. That's why Red Hat we are collaborating with Intel, AMD, and Arm. Together we're working upstream to create an any CPU enablement story directly within vLLM. By working in the open with the CPU vendors, we're

directly extending your accelerator inference choice to include your existing processors. We also understand that performance matters. That's why alongside our any CPU enablement story, we're also sharing vLLM CPU perf eval. It's an open source project that wraps automation around an LLM benchmarking tool called Guide LLM. And while we've built a solid foundation that benchmarks the key performance metrics that people are interested in like throughput and latency,

um to extend the the project or the framework, we really need community support. So what we really want is your feedback, your use cases, your PRs, your issues, all in the hope of standardizing the way CPU inference benchmarking is done today. So, if that's something that you're interested in getting involved with, or you would like to see a demo of VLM CPU perf eval, then please drop

by the Red Hat booth later. Thank you very much.