PyTorch Conference Europe 2026

Optimizing Large MoE Inference on NVIDIA Blackwell: NVFP4, ADP, and DualPipe Strat... Julien Demouth

21:35 · 07 Apr 2026 – 08 Apr 2026 · YouTube

About this talk

This talk focuses on optimizing large Mixture-of-Experts (MoE) inference using NVIDIA Blackwell’s fifth-generation Tensor Cores, specifically for models like DeepSeek-V3/R1. The speaker discusses the transition to NVFP4 precision for MoE weights to lower memory requirements, along with FP4/FP8 key-value caching to enhance concurrency and reduce the attention layer's memory footprint. The session explores Expert Parallelism for optimizing FLOPS in expert layers and Attention Data Parallelism to prevent redundant key-value replication, transforming Multi-Head Latent Attention into Multi-Query Attention. Advanced execution strategies are also addressed, such as DualPipe algorithms for efficient computation and communication, integration of DeepGEMM and FlashInfer kernels, as well as runtime optimizations through Programmatic Dependent Launch and CUDA Graphs to diminish host latency and improve speculative decoding with Multi-Token Prediction.