Scraping Flower Metrics with Prometheus
Flower is best known as Celery's web dashboard, but since version 1.0 it also exposes a Prometheus endpoint. Because Flower already listens to Celery's event stream, pointing Prometheus at it is the quickest way to get task counts, runtimes, and worker status into your existing monitoring. This guide sets it up and explains where it falls short, as part of Flower for Celery Monitoring in Observability & Monitoring for Job Queues.
Problem Statement
A team runs Celery with RabbitMQ and uses Flower to look at tasks when something seems wrong. They have no alerting on Celery at all; an outage last month — workers lost their broker connection and stopped consuming for 40 minutes — was reported by a customer. Prometheus and Grafana already monitor the web tier. They want task success and failure rates, runtime percentiles, and an alert when workers go offline, and they would rather not add a new component if Flower can provide the data.
Prerequisites
- Flower 1.2 or newer running against the same broker as the workers.
- Celery workers started with events enabled (
-Eorworker_send_task_events = True). - Prometheus able to reach Flower over the network.
- Flower's authentication configured, since the metrics endpoint sits behind it by default — see securing the Flower dashboard in production.
Step 1 — Turn On Task Events
Flower builds all of its knowledge from Celery events: task-received, task-started, task-succeeded, task-failed, task-retried, and worker heartbeats. Workers do not send task events by default, so without them Flower sees only heartbeats and its task metrics stay at zero.
# celeryconfig.py
worker_send_task_events = True # same as starting workers with -E
task_send_sent_event = True # adds task-sent from producers, for queue-time metrics
Events add a small amount of broker traffic — one message per state change per task. At thousands of tasks per second that becomes significant; Step 6 covers the trade-off.
Step 2 — Check the Endpoint
Flower serves metrics at /metrics on its normal port. Confirm it before configuring Prometheus:
curl -s -u "$FLOWER_USER:$FLOWER_PASS" http://flower:5555/metrics | grep '^flower_' | head
# flower_events_total{task="billing.charge",type="task-succeeded",worker="celery@w1"} 1832.0
# flower_task_runtime_seconds_bucket{le="0.5",task="billing.charge",worker="celery@w1"} 1790.0
# flower_worker_online{worker="celery@w1"} 1.0
# flower_worker_number_of_currently_executing_tasks{worker="celery@w1"} 3.0
# flower_task_prefetch_time_seconds{task="billing.charge",worker="celery@w1"} 0.004
The metrics that matter most:
| Metric | Type | Meaning |
|---|---|---|
flower_events_total |
counter | events by task name, event type, and worker |
flower_task_runtime_seconds |
histogram | run time of succeeded tasks |
flower_worker_online |
gauge | 1 while a worker's heartbeats are arriving |
flower_worker_number_of_currently_executing_tasks |
gauge | tasks in progress per worker |
flower_task_prefetch_time_seconds |
gauge | time between a task being received and started, per task and worker |
Step 3 — Configure the Scrape
Flower's authentication protects /metrics like every other page. Either give Prometheus its own credentials, or put a reverse proxy in front of Flower that allows /metrics without login only from Prometheus's network address and requires authentication for everything else. Do not turn authentication off entirely to make scraping easier: Flower's UI and API can revoke tasks and shut down workers. With basic auth:
scrape_configs:
- job_name: flower
scrape_interval: 30s
metrics_path: /metrics
basic_auth:
username: prometheus
password_file: /etc/prometheus/secrets/flower-password
static_configs:
- targets: ["flower.celery.svc:5555"]
If Flower runs with a URL prefix (--url_prefix=flower), the path becomes /flower/metrics. Run exactly one Flower instance per broker for metrics; two instances each counting the same events would double every rate in a dashboard that sums across targets.
Step 4 — Build Queries for Rates, Failures, and Runtime
Because flower_events_total counts every event type, most useful queries filter on type:
# tasks succeeded per second, by task
sum by (task) (rate(flower_events_total{type="task-succeeded"}[5m]))
# failure ratio per task
sum by (task) (rate(flower_events_total{type="task-failed"}[10m]))
/ sum by (task) (rate(flower_events_total{type=~"task-succeeded|task-failed"}[10m]))
# p95 runtime per task
histogram_quantile(0.95, sum by (task, le) (rate(flower_task_runtime_seconds_bucket[10m])))
# workers online
sum(flower_worker_online)
Step 5 — Add Alerts
Three alerts would have caught the 40-minute outage and the failure spikes that preceded it:
- alert: CeleryWorkersOffline
expr: sum(flower_worker_online) < 2
for: 3m
- alert: CeleryNoTasksSucceeding
expr: sum(rate(flower_events_total{type="task-succeeded"}[10m])) == 0
for: 10m
- alert: CeleryTaskFailureRatioHigh
expr: |
sum by (task) (rate(flower_events_total{type="task-failed"}[15m]))
/ sum by (task) (rate(flower_events_total{type=~"task-succeeded|task-failed"}[15m])) > 0.1
for: 15m
Workers that lose their broker connection also stop sending heartbeats, so flower_worker_online drops within the heartbeat timeout — about two minutes by default. Alert on "no successes" as well, because a worker can be online yet stuck.
Step 6 — Know the Limits and Fill the Gaps
Flower is a convenient metrics source, not a complete one:
- No queue depth. Flower does not read queue lengths from the broker, and backlog is the most important capacity signal. Add a broker exporter (RabbitMQ's built-in Prometheus plugin, or a Redis exporter reading list lengths) or instrument Celery with a Prometheus exporter.
- State is in memory. When Flower restarts, counters reset (Prometheus rates cope) and events sent while it was down are never counted.
- One process, all events. At very high task rates, a single Flower process can fall behind consuming events, delaying and distorting metrics. Beyond a few thousand tasks per second, a dedicated exporter or worker-side instrumentation scales better.
- Label cardinality. Metrics are labelled by worker hostname; with autoscaled workers whose names change on every deploy, series accumulate. Use
metric_relabel_configsto drop theworkerlabel where you do not need it.
Also note that Flower only sees events from workers connected to the broker it watches; with several Celery apps or virtual hosts, run one Flower per broker and label targets accordingly. Treat Flower's metrics as a fast start and a good source for worker online status, and plan to add broker-level queue metrics before relying on it for capacity alerts.
Step 7 — Combine Flower with a Broker Exporter
The most complete low-effort setup pairs Flower's event metrics with the broker's own metrics. Flower tells you what workers are doing; the broker tells you what is waiting for them. Together they answer every first question during an incident.
For RabbitMQ, enable the rabbitmq_prometheus plugin and scrape port 15692; the rabbitmq_queue_messages_ready and rabbitmq_queue_consumers metrics are the ones to graph. For Redis, the redis_exporter's check-keys option can export the lengths of the Celery queue lists. Alert on a queue that has messages ready but zero consumers — that single rule would have caught the outage in the problem statement within minutes, even if Flower itself had been down. Keep task names consistent with queue names where you can, so panels from both sources line up without complex joins.
Verification
flower_events_total{type="task-succeeded"}increases in Prometheus as tasks run.sum(flower_worker_online)matches the number of running workers, and drops within about two minutes after stopping one.- Stopping all workers in staging fires
CeleryWorkersOfflineand thenCeleryNoTasksSucceeding. - The Grafana panel for p95 runtime per task roughly matches runtimes seen in Flower's task view.
Gotchas & Edge Cases
Events off after a config change. A deploy that drops -E or worker_send_task_events silently turns task metrics to zero while worker metrics look normal. The "no tasks succeeding" alert catches this.
Remote control disabled. If workers run with worker_enable_remote_control = False, Flower cannot inspect them but can still count events; the metrics keep working.
Worker names with process IDs. Workers started without an explicit -n name get a hostname-based default, and in containers the hostname is the pod name, which changes on every deploy. Each new name creates new series in Prometheus while the old ones go stale. Name workers by role (-n billing@%h) and drop or aggregate the worker label in recording rules so dashboards stay readable across deploys.
Clock and heartbeat interval. flower_worker_online depends on heartbeats arriving within Flower's expected interval. Workers under heavy CPU load or with long blocking tasks in the solo pool can miss heartbeats and flap between online and offline. Use for: durations on the alert, and avoid the solo pool for production workers.
Timeouts on scrape. Flower serves metrics from the same Tornado process as its UI. A heavy UI session can slow scrapes; set a scrape timeout above the default 10 seconds if you see gaps.
FAQ
Can I run Flower only for metrics, without anyone using the UI? Yes. Many teams run Flower internally just as an event-to-metrics bridge, with the UI reachable only through a port-forward.
Does Flower's runtime histogram include failed tasks? It is fed from succeeded tasks. Failure timing appears only in event counts, so if slow failures matter to you, instrument that in the worker.
Should I use Flower or celery-exporter? Both consume events. Flower gives you a UI and metrics in one process; a dedicated exporter usually adds queue length and is lighter. Pick one as the metrics source so counts are not duplicated.
How long does Flower keep task history?
Flower keeps a bounded number of tasks in memory (--max_tasks, 10,000 by default) for its UI. That limit does not affect the Prometheus counters, which count every event since Flower started.
Related
- Flower for Celery Monitoring — what Flower shows and how to run it.
- Instrumenting Celery with a Prometheus Exporter — adding queue depth and more.
- Building a Celery Grafana Dashboard — panels built on these metrics.
- Running Flower on Kubernetes — deploying the Flower instance Prometheus scrapes.