About this talk
This talk by Luca Wehrstedt from Meta explores the evolution of training with low-precision float8 data types using NVIDIA's Hopper to Blackwell generations of GPUs. The speaker discusses how TensorCore acceleration revolutionizes training efficiency while navigating crucial decisions related to accuracy versus efficiency and precision versus range. The session highlights advancements such as the DeepSeek release and micro-scaling formats in Blackwell, providing a comprehensive comparison of these approaches to assist researchers in optimizing their training strategies.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17