Lightning Talk: Accelerating PyTorch Models With Torch.compile's C++ Wrapper Mode - Bin Bao, Meta
About this talk
This talk introduces torch.compile's C++ wrapper mode, a feature that significantly reduces CPU overhead and enhances the performance of PyTorch models. The speaker discusses the challenges posed by CPU overhead as GPU capabilities expand and compiler optimizations improve GPU kernel speed, presenting cpp-wrapper mode as a solution that generates optimized C++ code to address these issues. The session highlights how cpp-wrapper mode outperforms alternatives like CUDAGraphs, particularly in scenarios with dynamic input shapes, demonstrating a 39% speedup in benchmark results from the OSS Huggingface suite. Attendees will gain insights into effectively utilizing cpp-wrapper mode to overcome CPU-bound limitations while optimizing their machine learning applications.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17