PyTorch Conference Europe 2026

Enabling State-of-the-art Asynchronous Execution in Torch.compile With CUDA Streams - Michael Lazos

28:38 · 07 Apr 2026 – 08 Apr 2026 · YouTube

About this talk

This talk by Michael Lazos from Meta explores the enhancement of asynchronous execution in Torch.compile through the use of CUDA streams. It addresses how CUDA streams enable parallelization of GPU computation on NVIDIA GPUs, facilitating overlapping communication and compute kernels, as well as training on multiple batches in parallel. The speaker explains the concept of activation offloading to optimize memory usage and prevent out-of-memory errors by storing activations in CPU memory until they are needed. By integrating seamless CUDA stream support in PyTorch 2, Lazos demonstrates how users can utilize familiar eager APIs for stream assignment and synchronization directly within torch.compile, streamlining the workflow while improving efficiency for models with custom streaming patterns.