Measuring Queue Wait Time with Enqueue Timestamps
Queue length tells you how much work is waiting; it does not tell you how long anyone has been waiting. A backlog of 50,000 jobs that drains in two minutes is fine, and a backlog of 20 jobs that has been stuck for an hour is not. Wait time — the gap between when a job became ready and when a worker started it — measures what users actually experience. This guide adds it to any queue technology, as part of Prometheus Metrics for Workers in Observability & Monitoring for Job Queues.
Problem Statement
A team alerts when any queue has more than 10,000 jobs waiting. The alert fires every night during a scheduled batch that is designed to build a large backlog and drain it by morning, so on-call engineers have learned to ignore it. Meanwhile, a low-volume queue for account verification emails had a single stuck consumer for three hours, never exceeded 200 jobs, and never alerted. Users could not sign up. The team wants a metric that reflects how long real work waits, one that works across Celery, BullMQ, and a Go service using SQS, and alerts that fire when users are affected and stay quiet when a big backlog is draining as planned.
Prerequisites
- Control over the producer code (to add a timestamp) and the worker code (to record the metric).
- A metrics library in each worker:
prometheus_client,prom-client, or the Go client. - Synchronised clocks (NTP or chrony) on producers and workers.
- A list of which jobs are intentionally delayed or scheduled.
Step 1 — Define Wait Time Precisely
Wait time should mean "how long a job was ready to run but was not running." That excludes two kinds of intended waiting:
- Scheduled delays. A job enqueued with a 10-minute delay is not late during those 10 minutes.
- Retry backoff. A job waiting 5 minutes before its third attempt is waiting on purpose.
So the metric is started_at − ready_at, where ready_at = enqueued_at + requested_delay, recorded on the first attempt. If you want to know whether retries are also delayed beyond their backoff, record that as a separate metric with its own label; mixing it in makes the main number meaningless.
Step 2 — Stamp Jobs at Enqueue
Some queues record an enqueue time for you: BullMQ stores job.timestamp, Sidekiq stores enqueued_at, SQS provides the SentTimestamp attribute. Celery and many custom queues do not, so add it to the message headers or payload yourself. Record it as Unix epoch seconds (or milliseconds) in UTC, and record the requested delay alongside so the worker does not need to know how each producer scheduled the job:
# Celery: stamp in a before_task_publish signal so every producer does it
import time
from celery.signals import before_task_publish
@before_task_publish.connect
def stamp(headers=None, **_):
headers.setdefault("enqueued_at", time.time())
eta = headers.get("eta")
headers.setdefault("ready_at", _parse_eta(eta) if eta else headers["enqueued_at"])
// Go producer with SQS: put the ready time in a message attribute
readyAt := time.Now().Add(delay).UnixMilli()
input.MessageAttributes = map[string]types.MessageAttributeValue{
"ready_at": {DataType: aws.String("Number"), StringValue: aws.String(strconv.FormatInt(readyAt, 10))},
}
Stamping in one shared place — a publish signal, client middleware, or a small enqueue wrapper — ensures every job gets the timestamp, including jobs enqueued from scripts and other jobs.
Step 3 — Record a Histogram in the Worker
When a worker starts a job, compute the wait and observe it in a histogram labelled by queue (and optionally job name). A histogram, not a gauge or summary, lets Prometheus compute percentiles across all worker processes:
from prometheus_client import Histogram
from celery.signals import task_prerun
WAIT = Histogram("job_wait_seconds", "Time from ready to start",
["queue"], buckets=[0.1, 0.5, 1, 2, 5, 10, 30, 60, 300, 900, 3600])
@task_prerun.connect
def observe_wait(task=None, **_):
req = task.request
if req.retries: # only first attempts
return
ready_at = (req.headers or {}).get("ready_at")
if ready_at is None:
return
queue = req.delivery_info.get("routing_key", "unknown")
WAIT.labels(queue).observe(max(0.0, time.time() - float(ready_at)))
Choose buckets that bracket your targets: if the target for a queue is 30 seconds, you need bucket boundaries around 30 so the percentile is accurate there. See choosing histogram buckets for job duration — the same principles apply to wait time. For BullMQ, the event-based version is in instrumenting BullMQ with prom-client.
Step 4 — Cover Jobs That Have Not Started Yet
A histogram observed at start time has a blind spot: if no worker picks anything up, nothing is observed and the metric goes quiet exactly when it matters most — the stuck-consumer case from the problem statement. Complement it with the age of the oldest waiting job, read from the queue at scrape time:
Most brokers expose this directly: Sidekiq's Queue#latency, SQS's ApproximateAgeOfOldestMessage in CloudWatch, BullMQ by reading the first waiting job's timestamp, Celery with Redis by peeking at the head of the list and reading its ready_at header. Export it as a gauge from a single collector, not from every worker.
Step 5 — Alert on Wait, Not on Length
With both metrics in place, replace the queue-length alert with alerts that match user impact:
- alert: QueueWaitAboveTarget
expr: |
histogram_quantile(0.95, sum by (queue, le) (rate(job_wait_seconds_bucket[10m])))
> on (queue) group_left queue_wait_target_seconds
for: 10m
- alert: QueueStalled
expr: queue_oldest_job_age_seconds > 3 * on (queue) group_left queue_wait_target_seconds
for: 5m
Store each queue's target as a recording rule or a static metric (queue_wait_target_seconds{queue="verification"} 30), so every queue gets its own threshold without duplicating rules. The nightly batch queue can have a target of six hours; the verification queue, 30 seconds. For burn-rate alerting on these targets, see burn-rate alerts for queue backlogs.
Step 6 — Deal with Clock Skew
The enqueue timestamp comes from the producer's clock and the start time from the worker's. If they differ by two seconds, every wait time is off by two seconds — invisible for a 30-minute target, misleading for a 1-second target. Keep clocks synchronised and monitor offset (node_timex_offset_seconds from node_exporter). Clamp negative values to zero, as in the code above, and count how often you clamp; a steady stream of negative waits is a sign that a host's clock has drifted. Where available, prefer a timestamp set by the broker itself (SQS SentTimestamp, Kafka's log-append time) since it comes from one clock.
Verification
- Enqueue a job with no delay while workers are idle: its wait is under a second.
- Enqueue a job with a 60-second delay: its wait is also under a second, not 60.
- Stop all workers for a queue: the oldest-job age gauge rises steadily and
QueueStalledfires after the expected time. - Run the nightly batch: the wait alert stays quiet while the batch drains within its target.
- The count of clamped negative waits stays near zero.
Gotchas & Edge Cases
Prefetch hides waiting. Workers that prefetch messages hold them in local buffers. A job can wait in a worker's buffer after leaving the broker, so measure at the moment execution starts, not when the message is received.
Priority queues. Averages across priorities hide starvation of low-priority jobs. Label the histogram by priority level if you use them.
Re-enqueued jobs. Jobs moved back from a dead-letter queue keep their original timestamp and would record enormous waits. Reset ready_at when replaying, or skip the observation for replays.
Very large backlogs by design. For queues whose purpose is to absorb load, wait time is still the right metric; the target is simply larger.
FAQ
Why not use the average wait? Averages hide the jobs that waited longest, which are the ones users complain about. Use p95 or p99, and alert on them.
Is queue length ever useful? Yes — for capacity planning, autoscaling signals, and memory budgets. It just should not be the primary user-facing alert. Forecasting backlog drain time combines length with throughput to predict when a backlog clears.
Should the wait histogram include the job's own duration? No. Keep wait and processing time separate; end-to-end latency is their sum, and you can compute it or record it as a third metric if users care about completion time.
How do I measure wait time for Kafka consumers? Consumer lag in messages is the usual Kafka metric, but it has the same weakness as queue length. Compute time lag instead: the difference between now and the timestamp of the last committed record per partition. Tools such as Burrow or kafka-lag-exporter can export it, and it maps directly onto the oldest-job age gauge above.
What about jobs that never start because they are cancelled? They produce no wait observation, which is correct: nobody waited for their result. If cancellations are common, count them separately so a drop in observations is not mistaken for a stall.
Related
- Prometheus Metrics for Workers — what to measure and why.
- Defining SLOs for Job Latency — turning wait time into objectives.
- Alerting on Queue Backlog with Prometheus — complete alert rules.
- Instrumenting BullMQ with prom-client — the BullMQ implementation.