Lightning Talk: Scaling Recommendation Systems To 2K GPUs and Beyond - Zain Huda, Meta
About this talk
This talk covers the advancements in recommendation system scalability at Meta, focusing on the implementation of 2D sparse parallelism. The speaker explains how this technology enables the scaling of sparse recommendation embedding tables from 1,000 to 8,000 GPUs, facilitating the largest Ads model training runs in a production environment. The session also delves into the technical challenges associated with handling sparse operations and shares insights on optimizing these systems to enhance performance while reducing memory usage at scale. Attendees will gain valuable lessons from designing high-performance systems that leverage extensive GPU resources effectively.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17