Lightning Talk: Bringing Google’s Colossus to PyTorch: Rapid Stor... Ankita Luthra & Trinadh Kotturu
About this talk
This talk examines the transition from compute to storage bottlenecks in PyTorch models as they scale to billions of parameters. The speakers introduce Rapid Storage, which implements Google’s Colossus stateful protocol to enhance data access through persistent gRPC streams, effectively bypassing REST APIs. They discuss how this architectural shift results in less than 1ms random read/write latency, 20 times faster data access, and a remarkable 6 TB/s aggregate throughput, significantly reducing tail latency for random I/O. The session also explores the integration of Rapid Storage with gcsfs and the broader fsspec ecosystem, ensuring high-performance I/O across various data frameworks like Dask and Ray, ultimately helping users maximize GPU efficiency and achieve linear scaling in cloud environments.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17