Lightning Talk: Beyond Generic Spans: Distributed Tracing for Actio... Sally O'Malley & Greg Pereira
About this talk
This talk discusses the critical need for end-to-end observability in production large language models (LLMs) to monitor performance, attribute costs, and validate optimizations. The speakers detail their implementation of tracing for llm-d, a high-performance distributed LLM inference framework, using manual OpenTelemetry instrumentation to create actionable insights. Key topics include how distributed tracing can be utilized to track cache-aware routing, analyze processing choices, profile models across multi-node deployments, and correlate request patterns with workload autoscaling decisions. Attendees will gain an understanding of how LLMOps necessitates a fresh approach to distributed tracing compared to traditional microservices and learn effective strategies for instrumenting inference stacks.
More from this event
See all 103 talks →
What PyTorch Conference Europe 2026 Was Really Like – Official PyTorchCon EU Highlights | Paris
0:53
Lightning Talk: How DeepInverse Is Solving Imaging in Science and H... Andrew Wang & Minh Hai Nguyen
9:50
Why WideEP Inference Needs Data-Parallel-Aware Scheduling - Maroon Ayoub & Tyler Michael Smith
25:37
Write Once, Run Everywhere with Pytorch Transformers - Pedro Cuenca, Hugging Face
19:17