Lightning Talk: Step-Aligned Telemetry for Distributed PyTorch Training (Time... Abhinav Srivastav
About this talk
This talk covers the challenges of distributed PyTorch training, highlighting how misalignment between time-sampled telemetry and step-based training can lead to issues like throughput degradation and GPU idleness. The speaker, Abhinav Srivastav from TraceOpt, breaks down common failure modes in Distributed Data Parallel (DDP) training that are often overlooked by standard metrics, such as rank stragglers and step-time variance. This session demonstrates how step-aligned, rank-aware telemetry aggregation enhances debugging by providing clearer insights into performance issues and relating time and memory usage directly to training semantics without the need for complex profilers.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17