Lightning Talk: Distributed AI Without the Infrastructure Tax - Yahav Biran & Maen Suleiman
About this talk
This talk addresses the challenges of running distributed AI workloads in production, focusing on package compatibility, hardware abstraction, and network configuration. The speakers, Yahav Biran from Annapurna Labs and Maen Suleiman from Amazon, explain how AWS Neuron Deep Learning Containers (DLCs) resolve these issues by providing open-source, production-ready images for Trainium and Inferentia. They cover how DLCs combat dependency issues through versioning components like PyTorch and the Neuron SDK, utilize Dynamic Resource Allocation in Kubernetes to simplify hardware complexity, and ensure efficient data movement with pre-configured EFA driver settings. Attendees will learn strategies for scaling their applications from a laptop to a 32-node cluster while achieving significant resource utilization improvements and faster deployment times.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17