Lightning Talk: Achieving SOTA GEMM Performance: A CuTeDSL Backend for PyTorch Induc... Nikhil Patel
About this talk
This talk covers the development of a new CuTeDSL backend for PyTorch Inductor that aims to achieve state-of-the-art GEMM performance on NVIDIA's Blackwell architecture. The speaker, Nikhil Patel from Meta, discusses how existing Triton-based kernels struggle to adapt to fast-evolving hardware, resulting in users needing to create custom kernels. The new backend integrates NVIDIA’s kernel implementations directly into PyTorch's compilation framework, providing built-in support for various GEMM operations and allowing seamless updates as new architectural features emerge. Early results from vLLM inference and TorchTitan training are presented, highlighting the backend's ability to enhance GEMM performance while relieving developers of the burden of maintaining hand-written kernels.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17