Building a Celery Grafana Dashboard

A good Celery dashboard lets someone who has just been paged decide within a minute whether tasks are failing, waiting, or slow — and which task is responsible. Most Celery dashboards fall short of that because they graph whatever metrics are available rather than the questions that need answers. This guide builds one around those questions, as part of Grafana Dashboards for Queues in Observability & Monitoring for Job Queues.

Problem Statement

A Django company runs Celery with a Redis broker: 30 task types, four queues, and workers on Kubernetes. They have Prometheus metrics from a Celery exporter and an imported community dashboard with 40 panels. During the last incident, the on-call engineer scrolled through graphs of heap memory, event counts per worker, and prefetch times without finding the cause: one task type had started timing out against a slow partner API, filling the retry schedule and delaying everything else in its queue. You want a compact dashboard that shows the queue-level impact at the top, points to the responsible task below, and links to the deploy that caused it.

Prerequisites

  • Grafana 10 or newer with a Prometheus data source.
  • Celery metrics in Prometheus from one of: celery-exporter (danihodovic), Flower's /metrics, or worker-side instrumentation. See instrumenting Celery with a Prometheus exporter.
  • Queue length metrics for the broker (the exporter provides them for Redis and RabbitMQ).
  • Workers sending task events (-E), since exporters depend on them.

Step 1 — Pick One Metrics Source and Learn Its Names

Metric names differ between sources, so decide which one the dashboard uses and stay with it. This guide uses celery-exporter, whose main metrics are:

Metric Meaning
celery_task_sent_total tasks published (needs task_send_sent_event)
celery_task_received_total tasks received by a worker
celery_task_succeeded_total / celery_task_failed_total / celery_task_retried_total outcomes by task name
celery_task_runtime_bucket runtime histogram by task name
celery_queue_length messages waiting per queue
celery_worker_up 1 per worker sending heartbeats
celery_active_consumer_count consumers per queue

If you use Flower's metrics instead, replace them with flower_events_total{type=...} and flower_task_runtime_seconds — the panel structure stays the same. Whatever the source, run exactly one instance, or every rate will be multiplied by the number of exporters.

Choosing a Celery metrics source celery-exporter consumes events and also reads queue lengths from the broker, which makes it the most complete single source. Flower consumes the same events and adds a UI and control API, but does not report queue length. Worker-side instrumentation with prometheus_client records precise per-attempt timings in each process but requires code changes and still needs a separate source for queue length. Three sources, one choice per dashboard celery-exporter + task events + queue length + no code changes best single source Flower + task events + UI and control API − no queue length add a broker exporter in-worker metrics + precise timings + custom labels − code changes pair with queue length

Event-based sources have one weakness in common: at very high task rates, a single consumer of the event stream can fall behind, and during broker trouble the events themselves may be delayed. If your dashboard will be used to diagnose broker problems, make sure at least the queue-length and worker-up panels come from something that does not depend on events, such as the broker's own exporter. Mixing sources is fine as long as each panel is clear about where its numbers come from — put the source in the panel description.

Step 2 — Lay Out Rows from Impact to Cause

Arrange the dashboard so it is read top to bottom in the order an incident unfolds:

A Celery dashboard in four rows Row one holds four stat panels: total queue length, tasks succeeded per second, failure ratio, and workers up. Row two shows queue length and throughput per queue over time. Row three breaks down failures, retries and p95 runtime by task name, to point at the cause. Row four shows worker CPU, memory and restarts, collapsed by default. Read top to bottom: impact, where, why, resources queue length succeeded / s failure ratio workers up queue length per queue throughput per queue failures by task retries by task runtime p95 by task worker CPU, memory, restarts (collapsed row)

Four stat panels at the top give a one-glance answer: is work waiting, is it flowing, is it failing, are workers alive. Colour thresholds on the stats (green, amber, red) should match your alert thresholds so the dashboard and the pager agree. The second row shows where; the third, why; the fourth, whether the workers themselves are unhealthy.

Step 3 — Write the Queries

The core queries, with $queue and $task as template variables (Step 4):

# Row 1 — stats
sum(celery_queue_length{queue_name=~"$queue"})
sum(rate(celery_task_succeeded_total{name=~"$task"}[5m]))
sum(rate(celery_task_failed_total{name=~"$task"}[10m]))
  / (sum(rate(celery_task_succeeded_total{name=~"$task"}[10m])) + sum(rate(celery_task_failed_total{name=~"$task"}[10m])))
sum(celery_worker_up)

# Row 2 — per queue
sum by (queue_name) (celery_queue_length{queue_name=~"$queue"})
sum by (queue_name) (rate(celery_task_succeeded_total{queue_name=~"$queue"}[5m]))

# Row 3 — per task
topk(10, sum by (name) (rate(celery_task_failed_total{name=~"$task"}[10m])))
topk(10, sum by (name) (rate(celery_task_retried_total{name=~"$task"}[10m])))
histogram_quantile(0.95, sum by (name, le) (rate(celery_task_runtime_bucket{name=~"$task"}[10m])))

Use topk in the per-task panels so a dashboard with 30 task types shows only the ones that matter right now. Set the per-task time series legend to "table" mode sorted by maximum, which puts the worst task at the top of the legend. For runtime, a heatmap often reveals more than a single percentile line — see Grafana heatmaps for job duration.

Step 4 — Add Template Variables

Variables make one dashboard serve every queue and task without duplicating panels:

queue:  label_values(celery_queue_length, queue_name)        multi-value, include All
task:   label_values(celery_task_received_total, name)       multi-value, include All
env:    label_values(celery_worker_up, namespace)             single value

Chain them where it helps — for example, filter the task list by the selected queue if your metrics carry both labels. Keep the default at "All" so the dashboard opens on the whole system, and let alert links pre-select a specific queue (Step 6). Avoid a variable per worker: worker names change with every deploy on Kubernetes, and a long, shifting list is more confusing than helpful.

Step 5 — Annotate Deploys and Incidents

Many queue problems start at a deploy. Grafana annotations draw a vertical line on every panel at the moment of a deploy, which makes the connection obvious:

A deploy annotation on the retries panel On the retries-by-task panel, most task types stay near zero all afternoon. At 14:05 a deploy annotation appears as a vertical dashed line, and immediately afterwards retries for partner_sync climb from almost zero to about forty per second. The annotation text names the release, which points directly at the change to investigate or roll back. Retries by task, with a deploy annotation 13:30 14:05 14:40 deploy: release 2026.09.18-3 partner_sync all other tasks

Post an annotation from the deploy pipeline with Grafana's HTTP API, tagged so the dashboard can show it:

curl -s -X POST "$GRAFANA_URL/api/annotations" \
  -H "Authorization: Bearer $GRAFANA_TOKEN" -H "Content-Type: application/json" \
  -d "{\"tags\":[\"deploy\",\"celery-workers\"],\"text\":\"release $RELEASE\"}"

In the dashboard settings, add an annotation query filtered by the deploy tag. Add a second annotation source for incidents from your incident tool, so the history of past problems is visible on the same graphs.

Step 6 — Connect Alerts and Keep It Maintained

Each alert rule should link to the dashboard with its variables set, so the person paged lands on the right view. In Prometheus alert annotations, include a URL such as https://grafana.internal/d/celery?var-queue={{ $labels.queue_name }}&from=now-3h. In Grafana-managed alerts, set the dashboard and panel IDs on the rule.

Store the dashboard as JSON in the repository next to the worker code and provision it through Grafana's provisioning files or Terraform, so changes are reviewed and a lost Grafana instance can be rebuilt. Review it after each incident: remove panels nobody used, and add the one that would have shortened the investigation. The general principles are in Grafana dashboards for queues.

Verification

  • A newcomer can say within a minute whether tasks are waiting, failing, or slow, using only the top two rows.
  • Selecting a single queue in the variable updates every panel.
  • A test deploy creates an annotation visible on all panels.
  • An alert notification links to the dashboard with the right queue selected.
  • The dashboard JSON is in version control and provisioned automatically.

Gotchas & Edge Cases

Retries are not failures. Celery counts a retry separately from a failure. A task that retries for hours never appears in the failure ratio. Keep the retries panel next to failures.

Queue length with prefetch. Tasks prefetched by workers leave the broker queue and do not appear in celery_queue_length. With a high prefetch multiplier, the queue can look empty while hundreds of tasks wait in worker buffers. Graph reserved tasks per worker too, if your exporter provides them.

Rate windows vs scrape interval. A rate() window shorter than four scrape intervals produces gaps. With a 30-second scrape, use windows of at least two minutes.

Stat panels hide trends. A stat showing "12,400 waiting" says nothing about whether the number is rising or falling. Enable the sparkline on every stat panel, or pair each stat with a small time series, so the direction is visible at a glance.

Dashboards per environment. Staging and production usually share a dashboard with an environment variable. Make production the default, and colour the environment selector so nobody spends an incident looking at staging graphs.

Label names differ by exporter version. Some versions use queue_name, others queue. Check before copying queries from elsewhere.

FAQ

Should I start from a community dashboard? It is a useful source of queries, but trim it heavily. A dashboard that tries to show everything shows nothing clearly during an incident.

How many panels is too many? If the top row does not fit on one screen without scrolling, there are too many. Put detail in collapsed rows or linked drill-down dashboards.

Can the same layout work for other queue systems? Yes. The rows are the same for BullMQ, Sidekiq, or SQS; only the queries change. See building a BullMQ Grafana dashboard.

How far back should the default time range go? Three to six hours. That is long enough to see the start of most incidents and a deploy that caused them, and short enough that recent spikes are not flattened. Save wider ranges for capacity reviews.

Who should own the dashboard? The team that owns the workers. Platform teams can provide shared panels and conventions, but the task-level rows need knowledge of what each task does and which ones matter most.

Related