Orchestrating Thousands of GPUs: Engineering Patterns for Large-Scale Model Training - Krishnaswamy
About this talk
This talk explores the complexities involved in training large AI models, emphasizing the need for orchestration of multi-node GPU systems, effective communication, and thoughtful engineering trade-offs. It discusses the transition from traditional computing models to distributed training and provides insights into the reliability of such systems in production environments. The speaker covers real-world challenges in distributed data processing and introduces the five dimensions of parallelism—data, tensor, pipeline, expert, and context parallelism—along with practical heuristics for scaling AI training architectures across various hardware setups. The session also highlights the roles of gradient synchronization, collective operations, and fault tolerance, showcasing frameworks like NCCL, Gloo, and MPI within the context of distributed training systems.
More from this event
See all 126 talks →
AI Is Not the Risk. Architectural Drift Is - Sunil Kalkunte
17:39
Breaking the Monolith: Tesco’s Journey to Federated GraphQL with xAPI - Vishwas Chandrashekar
29:13
A Practical Introduction to LangChain4j - Venkat Subramaniam
1:01:28
Beyond the AI Models: How Lowe’s is Building the Store That Knows - Swaroop Shivaram
13:59