Enabling State-of-the-art Asynchronous Execution in Torch.compile With CUDA Streams - Michael Lazos
About this talk
This talk by Michael Lazos from Meta explores the enhancement of asynchronous execution in Torch.compile through the use of CUDA streams. It addresses how CUDA streams enable parallelization of GPU computation on NVIDIA GPUs, facilitating overlapping communication and compute kernels, as well as training on multiple batches in parallel. The speaker explains the concept of activation offloading to optimize memory usage and prevent out-of-memory errors by storing activations in CPU memory until they are needed. By integrating seamless CUDA stream support in PyTorch 2, Lazos demonstrates how users can utilize familiar eager APIs for stream assignment and synchronization directly within torch.compile, streamlining the workflow while improving efficiency for models with custom streaming patterns.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17