Lightning Talk: FlexAttention + FlashAttention-4: Fast and Flexible - Driss Guessous, Meta
About this talk
This talk features Driss Guessous from Meta, who discusses the integration of FlexAttention with FlashAttention-4 to enhance performance in attention research. FlexAttention has allowed researchers to prototype custom attention variants in PyTorch but has faced throughput challenges compared to FlashAttention-3. The speaker details the innovative use of CuTeDSL to optimize performance on Blackwell GPUs, resulting in significant speed improvements across various workloads. Key topics include how PyTorch's Inductor generates modifications for FlashAttention-4 and practical advice for users looking to adopt this new backend.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17