About this talk
Marty (Grafana Tempo maintainer) and Marco (Grafana Mimir maintainer) share three real incidents from Grafana Cloud and the architectural changes they made — one of the most transparent production deep-dives at GrafanaCON.
Incident 1 — Mimir "Query of Death" (March 2023): A customer's 40KB PromQL regex queries maxed out ingester CPUs and took down the largest Mimir cluster. Blocking the queries didn't stop it because the Go regex engine kept running after cancellation — a four-line loop nobody thought could take 15 minutes. Fixes shipped: regex unrolling (97% of Grafana Cloud regexes now skip the engine entirely), a cost-based query planner, and ingester overload protection. The deeper architectural problem — that queries could take down ingestion — led to the Kafka-based read/write path decoupling in Mimir 3.0 (using WarpStream in Grafana Cloud).
Incident 2 — Tempo Parquet Dictionary Bloat: Queriers were OOMing on trace lookups for a "tiny" 500-span trace. The trace had trickled in just a few spans per day over seven days with large high-cardinality JSON attached, creating massive dictionaries across every Parquet block. Reading one small trace required unpacking gigabytes of dictionaries in memory. Tempo 3.0 introduces per-attribute dictionary control, achieving 95% memory reduction on affected lookups.
Incident 3 — Mimir Queue Starvation: A small number of slow store-gateway queries gradually occupied all query workers, causing fast ingester queries to sit idle in the FIFO queue. Fix: a multidimensional query scheduler with separate lanes per execution path. Store gateway latency was also addressed by replacing Go memory-mapped I/O (which blocks the entire Go processor on page faults) with direct syscalls, and by rebuilding the Mimir query engine as fully streaming — significantly reducing memory usage and latency for long time-range and high-cardinality queries.
0:00 Introduction
1:28 Incident 1: The "Query of Death" in Mimir
4:59 Fixes: Regex Unrolling, Cost Planner, Overload Protection
7:24 Architectural Fix: Kafka-Based Decoupling in Mimir 3.0
8:15 Incident 2: Tiny Traces Causing Tempo OOMs
10:50 Fix: Attribute Limits and Tempo 3.0 Block Format
12:36 Incident 3: Slow Queries Blocking Fast Queries
15:26 Fix: Multidimensional Query Scheduler
17:54 Fix: Streaming Query Engine and Direct Disk I/O
19:12 Lessons: Design for Isolation and On-Call Culture
Links/resources:
Learn about Grafana Mimir: https://grafana.com/oss/mimir/?src=yt
Read about Grafana Tempo: https://grafana.com/oss/tempo/?src=yt
Get started with the Grafana Cloud forever-free tier: https://grafana.com/g/cloud
Have a question? Ask Grot, your AI helper: https://grafana.com/grot/
Reach out in our community forums: https://gra.fan/communityyf
---
Thanks for watching!
👍 Was this video helpful? Like and subscribe to our channel for more videos.
Connect with Grafana Labs:
X: (https://www.twitter.com/grafana)
LinkedIn: (https://www.linkedin.com/company/grafana-labs/)
Facebook: (https://www.facebook.com/grafana)
#Grafana #Observability #Mimir #Tempo #Kafka #SRE
More from this event
See all 24 talks →
Grafana 13 Deep Dive: Suggested Dashboards, Dynamic Dashboards, SQL Expressions, and More!
1:23:01
Real-time ML Inside a $2M Electron Microscope for Nuclear Powered Data Centers | Theia Scientific
23:05
When Observability Meets a Tamagotchi
5:57
How to Speed Up Grafana Dashboards 100x with ASAPQuery
10:21