PyTorch Conference Europe 2026

Lightning Talk: Achieving SOTA GEMM Performance: A CuTeDSL Backend for PyTorch Induc... Nikhil Patel

7:45 · 07 Apr 2026 – 08 Apr 2026 · YouTube

About this talk

This talk covers the development of a new CuTeDSL backend for PyTorch Inductor that aims to achieve state-of-the-art GEMM performance on NVIDIA's Blackwell architecture. The speaker, Nikhil Patel from Meta, discusses how existing Triton-based kernels struggle to adapt to fast-evolving hardware, resulting in users needing to create custom kernels. The new backend integrates NVIDIA’s kernel implementations directly into PyTorch's compilation framework, providing built-in support for various GEMM operations and allowing seamless updates as new architectural features emerge. Early results from vLLM inference and TorchTitan training are presented, highlighting the backend's ability to enhance GEMM performance while relieving developers of the burden of maintaining hand-written kernels.