Lightning Talk: Cross-Region Model Serving: PyTorch Inference, Observability... Suraj Muraleedharan
About this talk
This talk covers the challenges organizations face in deploying, monitoring, and operating PyTorch models for inference at scale across multiple regions. The speaker presents production-tested architectures for multi-region PyTorch inference and LLMOps workflows, addressing key topics such as latency-based routing, blue-green deployments, and automated failover. The session also explores observability methods using OpenTelemetry for distributed tracing and Prometheus/Grafana dashboards to track critical metrics. Attendees will gain insights into CI/CD pipelines for cross-region model deployment that include automated rollback and drift detection, along with practical serving architectures and dashboards utilizing open-source tools.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17