Seamless Integration: Custom Kernels in the Torch.compile Stack Wi... Kshiteej K, Masaki K & Pawel G
About this talk
This talk focuses on the integration of custom kernels within the Torch.compile stack to enhance high-performance PyTorch workflows. The speakers, Kshiteej Kalambarkar, Masaki Kozuki, and Pawel Gadzinski from NVIDIA, address the challenges of graph-breaks caused by custom operations that can hinder performance gains. They provide a roadmap for making extensions "compiler-aware," using the Transformer Engine project as a case study to demonstrate the custom_op extension point. Key topics include identifying and profiling graph-breaks, the registration process for custom operations, and strategies for managing complex logic that impacts graph capture, with a real-world emphasis on maintaining throughput. This session is tailored for developers interested in building custom PyTorch extensions within a compiled framework.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17