Optimizing Large MoE Inference on NVIDIA Blackwell: NVFP4, ADP, and DualPipe Strat... Julien Demouth
About this talk
This talk focuses on optimizing large Mixture-of-Experts (MoE) inference using NVIDIA Blackwell’s fifth-generation Tensor Cores, specifically for models like DeepSeek-V3/R1. The speaker discusses the transition to NVFP4 precision for MoE weights to lower memory requirements, along with FP4/FP8 key-value caching to enhance concurrency and reduce the attention layer's memory footprint. The session explores Expert Parallelism for optimizing FLOPS in expert layers and Attention Data Parallelism to prevent redundant key-value replication, transforming Multi-Head Latent Attention into Multi-Query Attention. Advanced execution strategies are also addressed, such as DualPipe algorithms for efficient computation and communication, integration of DeepGEMM and FlashInfer kernels, as well as runtime optimizations through Programmatic Dependent Launch and CUDA Graphs to diminish host latency and improve speculative decoding with Multi-Token Prediction.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17