All talks from PyTorch Conference Europe 2026
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
This talk reflects on PyTorch Conference Europe 2026, which took place on April 7–8 in Paris, France...
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
This talk by Andrew Wang from DeepInverse and Minh Hai Nguyen from Université de Toulouse explores h...
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
This talk, presented by Maroon Ayoub from IBM and Tyler Michael Smith from Red Hat, explores the cha...
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
This talk covers the integration of the Hugging Face transformers library, built on pure PyTorch, in...
Build PyTorch to Understand PyTorch - Vijay Janapa Reddi & Andrea Mattia Garavagno
This talk features Vijay Janapa Reddi from Harvard University and Andrea Mattia Garavagno from the U...
Lightning Talk: Why Your Forecasting Transformer Isn’t Working (And How To Fix It... Rosheen Naeem
This talk focuses on the challenges and solutions for improving forecasting transformers in Python,...
Lightning Talk: Bayesian Neural Networks With Variational Inference in PyTorch - Lars Heyen
This talk covers the implementation of Bayesian Neural Networks (BNNs) using Variational Inference i...
Lightning Talk: TerraKit: Standardising AI-Ready Geospatial Data... Rosie Lickorish & Romeo Kienzler
This talk introduces TerraKit, an open-source Python library designed to standardize the preparation...
Lightning Talk: ExecuTorch on Microcontrollers: Deploying PyTorch To... RJ Ascani & Matthias Cremon
This talk features RJ Ascani and Matthias Cremon from Meta as they discuss ExecuTorch, an extension...
Enabling State-of-the-art Asynchronous Execution in Torch.compile With CUDA Streams - Michael Lazos
This talk by Michael Lazos from Meta explores the enhancement of asynchronous execution in Torch.com...
The Token Slice: Implementing Preemptive Scheduling Via Chunked Decod... Maroon Ayoub & Kellen Swain
In this talk, Maroon Ayoub from IBM and Kellen Swain from Google present the concept of Chunked Deco...
Lightning Talk: Deep Learning in the Wild: Embedded PyTorch for... Taraqur Rahman & Owen O'Donnell
This talk focuses on the use of deep learning in wildlife conservation, specifically through the dep...
The Science and Practice of Open and Scalable LLM Evaluations - Grzegorz Chlebus, NVIDIA
This talk by Grzegorz Chlebus from NVIDIA delves into the science and practice of open and scalable...
Lights, Camera, Inference! Video Generation as a Service With VLLM-O... Ricardo Noriega & Doug Smith
This talk covers the exploration of video generation as a service using vLLM-Omni and the LTX-2 open...
Keynote: PyTorch Updates - Edward Yang, Research Engineer, Meta
This keynote by Edward Yang, a Research Engineer at Meta, covers the latest updates in PyTorch, a po...
Tour De Force: LLM Inference Optimization From Simple To Sophisticated - Christin Pohl, Microsoft
This talk by Christin Pohl from Microsoft focuses on optimizing large language model (LLM) inference...
Lightning Talk: Bringing Google’s Colossus to PyTorch: Rapid Stor... Ankita Luthra & Trinadh Kotturu
This talk examines the transition from compute to storage bottlenecks in PyTorch models as they scal...
Helion 1.0: A High-Level DSL for Performance Portable Kernels - Oguz Ulgen, Meta
This talk features Oguz Ulgen from Meta discussing Helion 1.0, a high-level domain-specific language...
torch.compile and Diffusers: A Hands-On Guide to Peak Performance - Sayak Paul, Hugging Face
This talk covers the integration of torch.compile with the Diffusers library to enhance the performa...
Optimizing PyTorch on CPU-GPU Coherent Platforms - Matthias Jouanneaux, Nvidia
This talk covers the optimization of PyTorch on Nvidia's GB200 coherent platform, highlighting the b...
Lightning Talk: Jigsaw: Domain and Tensor Parallelism for High-Resolution Inp... Deifilia Kieckhefen
This talk explores Jigsaw, a PyTorch library designed to enhance domain and tensor parallelism for h...
Why Classic IAM Collapses for Agents: Rethinking IAM for Agentic Systems - Parul Singh, Red Hat
This talk, presented by Parul Singh from Red Hat, explores the limitations of traditional Identity a...
Securing Agentic AI With PyTorch: Threat Modeling & LLM Red Teaming in Practice - Valeri Milke
This talk focuses on securing agentic AI systems built with PyTorch, highlighting the unique securit...
Keynote: Co-Evolution: How the Open Source Intelligence Stack Compounds - Mark Collier
This keynote presentation by Mark Collier, Executive Director of the PyTorch Foundation and General...
Lightning Talk: Implementing Single-Dim Strategies With Sharding Validator - Anshul Sinha, Meta
This talk, presented by Anshul Sinha from Meta, explores the implementation of Single-Dim Strategies...
Keynote: Community Led Open Source RL - Joe Spisak
This keynote features Joe Spisak, VP of Product and Head of Open Source at Reflection AI, discussing...
Lightning Talk: Flexible Deployment of PyTorch Models on MCU-Class... Robert Kalmar & Martin Pavella
This talk covers the efficient deployment of PyTorch models on MCU-class devices using ExecuTorch, a...
Lightning Talk: TorchJD: Jacobian Descent in PyTorch - Pierre Quinton & Valérian Rey
This talk covers Jacobian descent (JD), an advanced optimization technique that extends gradient des...
Lightning Talk: Cross-Region Model Serving: PyTorch Inference, Observability... Suraj Muraleedharan
This talk covers the challenges organizations face in deploying, monitoring, and operating PyTorch m...
Lightning Talk: Coding Agents for Compiler Construction: Beyond the... Reza Rahimi & Stefan Krassin
This talk focuses on coding agents as engineering tools for compiler construction, specifically with...
Optimizing Reinforcement Learning at Trillion-Parameter Scale - Songlin Jiang
This talk covers the implementation and optimization of reinforcement learning at the trillion-param...
Sponsored Session: TorchTPU: Expanding TPU Programmabil... Kat Ko, Claudio Basile & Jana van Greunen
This talk covers TorchTPU, a new initiative by Google aimed at expanding TPU programmability to PyTo...
Lightning Talk: Training Embedding Model Resiliently for Multimodal M... Huamin Chen & Haichen Zhang
This talk explores the training of embedding and classification models for efficient routing in Larg...
Lightning Talk: Ethical, Privacy and Sustainability Considerations in PyTorch S... Paula Mesa Macias
This talk by Paula Mesa Macias from Pau&Company examines the ethical, privacy, and sustainability co...
Brevitas Quantization Library - Pablo Monteagudo Lago, AMD
This talk covers the Brevitas Quantization Library, an open-source PyTorch resource developed by AMD...
Sponsored Keynote: From One Node to Distributed Training and Inference. How the PyTo... Ramine Roane
This keynote presentation by Ramine Roane, Corporate Vice President of AI Product Management and Eco...
Teaching PyTorch To Read Your Worst PDFs With Docling - Mingxuan Zhao, Peter Staar & Carol Chen
This talk focuses on the challenges of extracting clean, structured data from real-world documents,...
Lightning Talk: Running ExecuTorch Applications With Silicon Accelera... George Gekov & Aki Makkonen
This talk explores the efficient deployment of machine learning models on low-power embedded systems...
Lightning Talk: From Pretrained To Personal: Privacy-First... Daniel Holanda Noronha & Iswarya Alex
This talk explores the advancements in using PyTorch on AI PCs for meaningful model fine-tuning whil...
Keynote: The Unbearable Lightness of (Agentic) Evaluations - Besmira Nushi
This talk by Besmira Nushi, Senior Manager of AI Research at NVIDIA, explores the evolving disciplin...
Parameterized CUDA Graph Launch in PyTorch: CUDA Graphs Without the Pain - Daniel Galvez, NVIDIA
This talk focuses on the challenges of using CUDA Graphs in PyTorch, highlighting the potential bott...
On-Device LLM Inference on Android With ExecuTorch and Qualcomm QNN - Shivay Lamba & Kartikey Rawat
This talk covers the execution of large multimodal models like CLIP on Android devices using ExecuTo...
Sponsored Keynote: Any [ Agent | Model | Accelerator | Cloud ]... Maryam Tahhan & Nicolò Lucchesi
This talk features Maryam Tahhan and Nicolò Lucchesi from Red Hat, who discuss the advancements in o...
Lightning Talk: Combo Kernels: Horizontal Fusion Optimizat... Karthick Panner Selvam & Elias Ellison
This talk by Karthick Panner Selvam and Elias Ellison from Meta focuses on combo kernels, a compiler...
How To Write C++ Extensions in 2026 - Jane Xu, Meta & Mikayla Gawarecki, Meta
This talk covers the creation of C++ custom operation extensions for PyTorch in 2026, presented by J...
Lightning Talk: Inside VLLM's KV Offloading Connector: Async Memory Transfers for... Nicolò Lucchesi
This talk covers the KV Offloading Connector in vLLM 0.11.0, which addresses the challenge of recomp...
Lightning Talk: Accelerating On-Device ML Inference With ExecuTorch and Arm SME2 - Jason Zhu, Arm
This talk explores the challenges of achieving low-latency inference for complex on-device AI worklo...
From Responses To Trajectories: Multi-Turn and Multi-Environ... Kashif Rasul & Sergio Paniego Blanco
This talk covers the advancements in post-training large language models (LLMs) using reinforcement...
Lightning Talk: From Hugging Face To Handheld: Scaling LLM Deployment W... Cormac Brick & Weiyi Wang
This session explores the end-to-end journey of deploying custom PyTorch-based Open Source Large Lan...
Beyond the Theory: What Actually Breaks When You Scale Your Disaggregat... Ekin Karabulut & Ron Kahn
This talk focuses on the challenges of scaling disaggregated PyTorch models for inference in high-de...
Bringing PyTorch Monarch to AMD GPUs: Single-Controller Distributed Tra... Liz Li & Zachary Streeter
This talk covers the integration of PyTorch Monarch with AMD GPU technologies, focusing on single-co...
Lightning Talk: Why Logging Isn’t Enough: Making PyTorch Training Regressions Vi... Sahana Venkatesh
This talk presents an innovative approach to identifying training regressions in PyTorch, emphasizin...
Lightning Talk: Graph Based Pipeline Parallelism - Sanket Purandare, Meta & Simon Fan, Meta PyTorch
This talk discusses the challenges of pipeline parallelism in large models and how current PyTorch i...
Lightning Talk: Ball Tracking and Detection in Soccer Videos - Comparison of... Maciej Szymkowski
This talk focuses on ball tracking and detection in soccer videos, comparing Vision-Language Models...
Lightning Talk: Backpropagation-Free Optimization in PyTorch - Andrii Krutsylo
This talk, presented by Andrii Krutsylo from the Polish Academy of Sciences, explores backpropagatio...
Lightning Talk: Every Millisecond Counts: The Fine-tuning Journey of an Ultra-Eff... Pavel Macenauer
This talk, presented by Pavel Macenauer from NXP Semiconductors, delves into the optimization journe...
Bringing ExecuTorch To the Next Frontiers of Edge AI - Mergen Nachin, Meta
This talk features Mergen Nachin from Meta discussing the ongoing advancements of ExecuTorch, an on-...
Lightning Talk: Full-Stack PyTorch Robotics VLA: From Data To... Samet Akcay & Dmitriy Pastushenkov
This talk covers the transition of Vision-Language-Action models to production in robotics, addressi...
Lightning Talk: Slash LLM Cold-Start Times by Pre-distributing GPU... Billy McFall & Maryam Tahhan
This talk focuses on optimizing Large Language Model (LLM) deployments by addressing GPU cold-start...
Lightning Talk: Not All Tokens Are Equal: Semantic KV-Cache for Agen... Maroon Ayoub & Hyunkyun Moon
This talk, presented by Maroon Ayoub from IBM Research and Hyunkyun Moon from moreh, explores the co...
Lightning Talk: Debugging the Undebuggable: Introducing Torch.distributed.debug - Tristan Rice
This talk covers the complexities of debugging in distributed training using PyTorch and introduces...
Lightning Talk: KV-Cache Centric Inference: Building a State-Aware... Maroon Ayoub & Martin Hickey
This talk focuses on KV-cache centric inference, highlighting the significance of state management i...
Lightning Talk: FlexAttention + FlashAttention-4: Fast and Flexible - Driss Guessous, Meta
This talk features Driss Guessous from Meta, who discusses the integration of FlexAttention with Fla...
Lightning Talk: Beyond Generic Spans: Distributed Tracing for Actio... Sally O'Malley & Greg Pereira
This talk discusses the critical need for end-to-end observability in production large language mode...
Lightning Talk: Step-Aligned Telemetry for Distributed PyTorch Training (Time... Abhinav Srivastav
This talk covers the challenges of distributed PyTorch training, highlighting how misalignment betwe...
Optimizing Large MoE Inference on NVIDIA Blackwell: NVFP4, ADP, and DualPipe Strat... Julien Demouth
This talk focuses on optimizing large Mixture-of-Experts (MoE) inference using NVIDIA Blackwell’s fi...
Lightning Talk: Live Migration of PyTorch GPU Nodes From Azure To European Clouds - Mike Krom
This talk covers the live migration of PyTorch GPU nodes from Azure to European cloud providers, add...
Lightning Talk: Enabling the Audio Modality for Language Models - Eustache Le Bihan, Hugging Face
This talk features Eustache Le Bihan from Hugging Face, who discusses the integration of audio into...
Optimizing CPU LLM Inference in PyTorch: Lessons From VLLM - Crefeda Rodrigues & Fadi Arafeh
This talk covers optimizing CPU-based large language model inference using vLLM as a case study in t...
TorchStore: What We Learned Building Distributed Storage Sol... Lucas P, Danielle P, Allen W, Amir A
This talk explores the challenges and solutions in building distributed storage systems for Asynchro...
Lightning Talk: Scaling Recommendation Systems To 2K GPUs and Beyond - Zain Huda, Meta
This talk covers the advancements in recommendation system scalability at Meta, focusing on the impl...
De-mystifying PyTorch for ASICs: When (and Why) To Move Your Development To AI A... Alpha Romer Coma
This talk explores the transition of machine learning development from traditional GPUs to AI accele...
Model-Changing Transforms With Torch.compile - Thomas Viehmann, Lightning AI
This talk covers the capabilities of torch.compile, a tool designed to enhance the performance of Py...
PyTorch Symmetric Memory + NCCL Device APIs: A New Path Towards Multi-GP... Ke Wen & Sylvain Jeaugey
This talk covers the advancements in multi-GPU kernel development using PyTorch Symmetric Memory and...
Lightning Talk: Building a PyTorch‑native VLLM Plugin for IBM Spyre - Thomas Parnell & Thomas Ortner
This talk covers the redesign of the vLLM plugin for IBM Spyre, an AI accelerator employed in IBM Z...
PyTorch on RISC-V: From Cross-Compilation To Native CI - Ludovic Henry, Meta
This talk covers the efforts to port PyTorch natively to the RISC-V architecture, which is becoming...
Sponsored Keynote: Open Source Infrastructure for the AI Native Era - Jonathan Bryce
This keynote by Jonathan Bryce, Executive Director of the Cloud Native Computing Foundation, explore...
Keynote: Gemma 4: Compacting Intelligence for the Edge - Léonard Hussenot
This talk explores the philosophy and engineering behind Gemma 4, emphasizing that the future of AI...
Keynote: Stream Everything - Moving from Request input to Streaming input - Patrick von Platen
This talk features Patrick von Platen, a Research Engineer at Mistral AI, who presents insights on t...
Lightning Talk: Pluggable PyTorch LLM Inference Architecture With VLL... Yahav Biran & Maen Suleiman
This talk covers the evolution of PyTorch-based large language model (LLM) serving, focusing on the...
Lightning Talk: Faster Than SOTA Kernels in Torch.compile With Subgrap... Elias Ellison & Paul Zhang
This talk discusses how subgraph optimization and custom operator autotuning in torch.compile can ac...
From Gradients To Governance: Making PyTorch Lineage-Aware - Kateryna Romashko & Clodagh Walsh
This talk, presented by Kateryna Romashko and Clodagh Walsh from Red Hat, explores the necessity of...
Lightning Talk: Torch-Spyre: Compiling To a Multi-core Dataflow Accelerator... D. Grove & O. Tardieu
This talk covers the development of Torch-Spyre, an open source project designed to compile for the...
Portable High‑Performance LLM Serving: A Triton Backend for... Burkhard Ringlein & Jan van Lunteren
This talk covers the introduction of a Triton backend for vLLM, which is becoming the industry stand...
Lightning Talk: Bridging the Gap: Engineering Compliant... Muhammad Saqib Hussain & Mohaddisa Maryam
This talk features Muhammad Saqib Hussain and Mohaddisa Maryam from Neurosonic, focusing on creating...
Keynote: The Hub as Infrastructure. From Open PyTorch Models, to a Safe and Perfor... Lysandre Debut
This keynote features Lysandre Debut, Chief Open-Source Officer at Hugging Face, who discusses the c...
Seamless Integration: Custom Kernels in the Torch.compile Stack Wi... Kshiteej K, Masaki K & Pawel G
This talk focuses on the integration of custom kernels within the Torch.compile stack to enhance hig...
Keynote: PyTorch CTO - Matt White, Global CTO of AI, Linux Foundation
In this keynote, Matt White, Global CTO of AI and CTO at PyTorch Foundation, delivers an insightful...
Deploying PyTorch Models To the Browser and Beyond With Transformers.js - Joshua Lochner
This talk covers the engineering roadmap for running Hugging Face Transformers locally in web browse...
Lightning Talk: Building AI That Ops Teams Actually Trust - Robert King
This talk focuses on building AI systems that operations teams can trust, addressing the skepticism...
Lightning Talk: Distributed AI Without the Infrastructure Tax - Yahav Biran & Maen Suleiman
This talk addresses the challenges of running distributed AI workloads in production, focusing on pa...
Fp8 Training From Hopper To Blackwell - Luca Wehrstedt, Meta
This talk by Luca Wehrstedt from Meta explores the evolution of training with low-precision float8 d...
DualPipe from Scratch: Implementing DeepSeek's 5D Parallelism in PyTorch - Dev Jadhav, ING Bank
This talk covers the implementation of DeepSeek's 5D parallelism and DualPipe in PyTorch, addressing...
Building Trust for Users and Regulators Alike: A Cost-Efficient PyTorch Pat... Raja Gopal Hari Vijay
This talk explores the concept of Compliance-as-Code, focusing on how it integrates regulatory contr...
Keynote: vLLM & Ray Updates - Tyler Michael Smith & Artur Niederfahrenhorst
This keynote session features Tyler Michael Smith, Chief Architect of Inference Engineering at Red H...
Beyond JSON-RPC: Scaling Model Context Protocols With gRPC in the Py... Ashesh Vidyut & Madhav Bissa
This talk by Ashesh Vidyut and Madhav Bissa from Google addresses the limitations of Model Context P...
Lightning Talk: Achieving SOTA GEMM Performance: A CuTeDSL Backend for PyTorch Induc... Nikhil Patel
This talk covers the development of a new CuTeDSL backend for PyTorch Inductor that aims to achieve...
Bridging the Hardware Gap With Code Harnesses on the Hugging Face Kernels Hub - Ben Burtenshaw
This talk covers the development of code harnesses for standardizing kernel writing in the context o...
Sponsored Session: Fault-Tolerant Training: How We Build Rel... Cyril Konkratenko & Maurits de Groot
This talk covers fault-tolerant training for large-scale distributed AI workloads, highlighting Nebi...
Accelerating Complex-Valued Tensors With Torch.compile - Hameer Abbasi, OpenTeams Inc.
This talk covers the advancements in supporting complex-valued tensors with Torch.compile, presented...
Lightning Talk: Accelerating PyTorch Models With Torch.compile's C++ Wrapper Mode - Bin Bao, Meta
This talk introduces torch.compile's C++ wrapper mode, a feature that significantly reduces CPU over...
Lightning Talk: Trinity Large - Torchtitan on 2000+ B300s - Matej Sirovatka, Prime Intellect
This talk covers the use of torchtitan for scaling the training of ultra-sparse mixture-of-experts m...
Lightning Talk: Monarch: An API To Your Supercomputer - Marius Eriksen, Meta
This talk covers Monarch, a distributed programming framework for PyTorch that simplifies supercompu...