Lightning Talk: Faster Than SOTA Kernels in Torch.compile With Subgrap... Elias Ellison & Paul Zhang
About this talk
This talk discusses how subgraph optimization and custom operator autotuning in torch.compile can achieve performance improvements that surpass state-of-the-art results for matrix multiplication and distributed collective operations. The speakers, Elias Ellison and Paul Zhang from Meta, explain the innovative DecomposeK optimization, which enhances matrix multiplication speed, particularly when the inner dimension is large, offering up to a 28% speedup with activation fusion. They also introduce the Custom Op Autotuning feature, which benchmarks and identifies the fastest kernel implementations for custom operations, alongside a range-based dispatch autotuning method that dynamically selects optimal implementations based on input shapes. The results from their demo showcase significant performance gains, outperforming existing kernels by notable margins.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17