Bringing PyTorch Monarch to AMD GPUs: Single-Controller Distributed Tra... Liz Li & Zachary Streeter
About this talk
This talk covers the integration of PyTorch Monarch with AMD GPU technologies, focusing on single-controller distributed training using ROCm. The speakers, Liz Li and Zachary Streeter from AMD, discuss the innovative distributed programming paradigm that Monarch offers, which allows for orchestration of GPU clusters via a single Python script. They detail the technical efforts involved in adapting Monarch's GPU runtime and distributed communication stack to ROCm, including the transition from CUDA-specific components, memory management, and advanced GPU-to-GPU communication. Additionally, the session shares insights gained from deploying Monarch on MI300-class clusters, addressing performance, debugging, and enhancements to the developer experience, while emphasizing the ease of scalability in heterogeneous hardware setups.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17