Capacity Planning for Job Queues

A queue hides capacity problems until they become outages: it absorbs a shortfall silently, the backlog grows, and by the time anyone notices, jobs are hours late. This guide turns worker and broker sizing into arithmetic you can check, as part of Observability & Monitoring for Job Queues. It uses a handful of quantities that every queue already exposes — arrival rate, service time, concurrency, and queue wait — and the queueing relationships that connect them.

The questions it answers are the ones that come up in every planning review: how many workers do we need for next quarter's traffic, how long will a backlog take to drain after an outage, where will the broker saturate, and how much headroom is enough without paying for idle machines all night.

The Scenario: Black Friday Sizing

An e-commerce platform's order-processing queue runs 24 worker pods at 8 concurrent jobs each. On a normal day it handles about 60 orders per second at peak with queue waits under two seconds. Marketing forecasts that Black Friday will bring four times the normal peak for about three hours, with a flash sale at 18:00 that could spike to ten times for fifteen minutes. The previous year the queue fell six hours behind and customers received confirmations the next morning. The team needs to decide how many workers to run, whether autoscaling will react fast enough, whether the Redis broker will hold, and what backlog to expect from the flash sale — before the day, not during it.

The load to plan for Arrival rate over the day. The normal peak is 60 orders per second. From 15:00 to 21:00 the rate plateaus at 240 per second, four times normal. At 18:00 a flash sale spikes it to 600 per second for fifteen minutes. The dashed line marks current capacity at about 190 jobs per second, below both the plateau and the spike. Forecast arrivals, jobs per second current capacity ~190/s flash sale 600/s, 15 min plateau 240/s 00:00 24:00

Architectural Overview: The Four Numbers

Every capacity question for a queue reduces to four measured quantities:

  • λ (arrival rate) — jobs enqueued per second. Measured from enqueue counters, forecast from business metrics.
  • S (service time) — seconds a worker slot spends on one job, from start to acknowledgement. Use a distribution, not an average: p50 for throughput, p95/p99 for tail behaviour.
  • c (concurrency) — total worker slots across the fleet: pods × processes × threads (or goroutines, or async concurrency).
  • W_q (queue wait) — time between enqueue and start. The user-facing symptom, and the quantity an SLO should target, as in defining SLOs for job latency.

Two relationships connect them. Capacity is c / S: 192 slots each finishing a job every 0.8 seconds process 240 jobs per second. Utilisation is ρ = λ × S / c: the fraction of slots busy. As ρ approaches 1, queue wait does not grow linearly — it explodes, because every burst above the average has nowhere to go. That non-linearity is why "we're only at 90% utilisation" is not reassuring for a queue.

# The core arithmetic, with the scenario's numbers
S_p50 = 0.8                      # seconds per order job (median)
slots = 24 * 8                   # pods x concurrency = 192
capacity = slots / S_p50         # 240 jobs/s theoretical maximum

for arrivals in (60, 190, 240, 600):
    rho = arrivals * S_p50 / slots
    print(f"lambda={arrivals:>3}/s  utilisation={rho:5.0%}  "
          f"{'stable' if rho < 1 else 'backlog grows by %d/s' % (arrivals - capacity)}")
# lambda= 60/s  utilisation=  25%  stable
# lambda=190/s  utilisation=  79%  stable
# lambda=240/s  utilisation= 100%  backlog grows by 0/s   (unstable in practice)
# lambda=600/s  utilisation= 250%  backlog grows by 360/s

The detailed treatment of the sizing formula, including how variance in service time changes the answer, is in calculating worker count with Little's Law.

Wait time explodes near full utilisation A curve of relative queue wait against utilisation. Wait is small and nearly flat up to about 70 percent, rises noticeably by 85 percent, and climbs steeply toward infinity as utilisation approaches 100 percent. The recommended planning band is 60 to 75 percent at forecast peak. Relative queue wait vs utilisation plan here 0% 60% 75% 100% past ~85%: bursts pile up

Implementation 1: Sizing for the Plateau

For sustained load, size so that forecast peak lands at 60–75% utilisation. The margin absorbs short bursts, variance in service time, and a worker pod being replaced, without queue wait climbing.

def slots_needed(arrivals_per_s: float, service_s: float, target_util: float = 0.7) -> int:
    return math.ceil(arrivals_per_s * service_s / target_util)

plateau = slots_needed(240, 0.8)          # 275 slots
pods = math.ceil(plateau / 8)             # 35 pods at concurrency 8
print(plateau, pods)                      # 275 35

Then check the dependencies those workers call. 35 pods × 8 slots = 280 concurrent jobs, each holding a database connection for most of its service time: the order database must accept 280 more connections, and whatever it does per job must scale to 240 per second. Capacity planning for workers is capacity planning for everything the workers touch — the database, third-party APIs with rate limits, and the broker itself. The concurrency-per-pod choice has its own trade-offs, covered in right-sizing worker concurrency per CPU.

Implementation 2: Planning for Bursts and Drain Time

Sizing for a 15-minute spike at 600 per second would mean 86 pods idling most of the year. The alternative is to let the queue absorb the burst and plan how long it takes to drain. That is exactly what queues are for — as long as the resulting delay is acceptable and known in advance.

def burst_backlog(arrivals: float, capacity: float, minutes: float) -> float:
    return max(0.0, (arrivals - capacity) * minutes * 60)

def drain_minutes(backlog: float, capacity: float, arrivals_after: float) -> float:
    spare = capacity - arrivals_after
    return float("inf") if spare <= 0 else backlog / spare / 60

cap = 35 * 8 / 0.8                                    # 350 jobs/s with 35 pods
backlog = burst_backlog(600, cap, 15)                 # 225,000 jobs queued by 18:15
print(round(drain_minutes(backlog, cap, 240)))        # ~34 minutes to drain at the plateau

With 35 pods, the flash sale leaves 225,000 jobs queued, drained about 34 minutes after the spike ends; the worst-case order waits roughly 50 minutes. If that is too long, the options are pre-scaling before 18:00, prioritising payment-confirmation jobs over, say, recommendation updates (see Priority Queues & Job Fairness), or accepting the delay with a customer-facing message. The forecasting technique for live incidents is in forecasting backlog drain time.

Absorb the spike, then drain Backlog over time around the flash sale with 35 pods. From 18:00 to 18:15 arrivals exceed capacity by 250 jobs per second, so the backlog rises to 225,000. After the spike, spare capacity is 110 jobs per second at the plateau arrival rate, so the backlog falls back to zero at about 18:49. Backlog with 35 pods 225,000 queued at 18:15 +250/s drain at 110/s 18:00 18:15 ~18:49

Trade-off Analysis: Provisioning Strategies

Strategy Cost Queue wait during spike Risk
Static for peak spike (86 pods all day) Highest Minimal Paying for idle capacity 99% of the time
Static for plateau (35 pods), absorb spike Moderate ~50 min worst case Acceptable only if the delay is
Reactive autoscaling on backlog Low Depends on scale-up speed Scale-up lag: pods take 1–3 min, a spike lasts 15
Scheduled pre-scaling + autoscaling Low to moderate Small Forecast error; needs a schedule owner
Priority split (critical jobs isolated) Moderate Minimal for critical, longer for rest Needs job classification

Reactive autoscaling alone is usually too slow for a sharp spike: KEDA or an HPA sees the backlog, requests pods, waits for scheduling, image pulls, and startup — typically one to three minutes, plus node provisioning if the cluster must grow. For known events, schedule the scale-up ahead of time and let autoscaling handle the error in the forecast. The mechanics are in scaling workers with KEDA on queue length.

Failure Modes & Recovery

Scaling workers past a dependency's limit. Adding pods increases pressure on the database or an external API. Past its limit, service time rises and capacity falls — more workers make the queue slower. Remediation: find the scarcest dependency, cap total concurrency there, and scale workers only up to that cap.

Broker saturation. Redis is single-threaded per shard; RabbitMQ queues are bound to one core each; SQS has per-queue API limits for FIFO. At high rates the broker becomes the ceiling regardless of worker count. Remediation: measure the broker's own ceiling in a load test, as in benchmarking Redis broker throughput, and shard queues before reaching it.

Service time that changes under load. Planning uses today's service time, but under load, cache hit rates fall and database contention rises, so S grows. Remediation: measure S at the target load in a load test rather than extrapolating from quiet periods.

Retry amplification. During a partial outage, failing jobs retry and add to λ. A 10% failure rate with three retries adds up to 30% more load exactly when capacity is impaired. Remediation: include retry load in the plan and use circuit breakers — see circuit breakers for worker dependencies.

Performance Tuning and Measurement

Capacity plans are only as good as the measurements behind them. Instrument these, per queue:

# λ: arrival rate
sum by (queue) (rate(jobs_enqueued_total[5m]))

# Capacity actually used: busy slots / total slots
sum by (queue) (worker_busy_slots) / sum by (queue) (worker_total_slots)

# S: service time percentiles
histogram_quantile(0.95, sum by (le, queue) (rate(job_duration_seconds_bucket[10m])))

# W_q: queue wait from enqueue timestamps
histogram_quantile(0.95, sum by (le, queue) (rate(job_queue_wait_seconds_bucket[10m])))

Queue wait needs an enqueue timestamp on each job and a histogram observed at start — measuring queue wait time with enqueue timestamps shows how. Record utilisation per pool, not only per fleet: a fleet at 55% can hide one pool at 95% whose queue is quietly growing while the others idle. With all four signals on one dashboard, a planning review can read utilisation at last peak directly and apply the growth forecast to it, instead of arguing from anecdotes.

Validating the Plan with a Load Test

Arithmetic gives the plan; a load test proves it. The single most valuable test before a known peak replays the forecast shape — plateau, spike, drain — against a staging environment with production-like data volume, and checks three numbers against the plan: the capacity actually reached, the service time at that load, and the drain time after the spike.

# loadtest-plan.yaml — the forecast, expressed as stages for the load generator
stages:
  - { name: warmup,   rate: 60,  duration: 10m }
  - { name: plateau,  rate: 240, duration: 30m }
  - { name: spike,    rate: 600, duration: 15m }
  - { name: recovery, rate: 240, duration: 45m }
assert:
  plateau.utilisation_p95: "<= 0.75"
  plateau.queue_wait_p95: "<= 5s"
  spike.backlog_max: "<= 250000"
  recovery.drain_minutes: "<= 40"
  dependencies.db_connections_max: "<= 400"

A plan that fails these assertions in staging is a plan that would have failed on the day; the difference is that in staging it costs an afternoon. The tooling, including generating realistic job mixes and measuring from the queue's own metrics rather than the generator's, is covered in load testing queue throughput.

Several Queues on Shared Workers

Real fleets rarely serve one queue. The order workers in the scenario also consume a notifications queue and a recommendations queue, and the capacity question becomes how to divide 280 slots among three workloads with different urgency. Summing the load is the first step: capacity is shared, so required slots are the sum of λ × S across queues, divided by the target utilisation.

queues = {                       # forecast plateau: arrivals/s, service seconds
    "orders":          (240, 0.8),
    "notifications":   (300, 0.1),
    "recommendations": (80,  1.5),
}
busy = sum(l * s for l, s in queues.values())      # 192 + 30 + 120 = 342 busy slots
print(math.ceil(busy / 0.7))                        # 489 slots at 70%: 62 pods, not 35

The recommendations queue alone needs 120 busy slots — more than a third of the fleet — for work that is useful but never urgent. Treating it as a separate workload changes the plan: run it on its own small worker pool (or shared workers with a low weight), let it fall behind during the spike, and size the order pool only for orders and notifications. Isolation also protects the critical path from a slow dependency in the non-critical one, the same reasoning as in preventing tenant starvation with weighted queues.

Split the non-critical load out At the forecast plateau, orders need 192 busy slots, notifications 30, and recommendations 120, for a total of 342, which requires about 489 slots at 70 percent utilisation. Moving recommendations to its own small pool that is allowed to fall behind during the spike leaves the critical pool sized for 222 busy slots. Busy slots at the plateau, by queue orders 192 notif 30 recommendations 120 critical pool: sized for 222 busy slots own pool, may lag Summing load across queues is the plan; splitting pools is how the plan survives a spike.

Writing the Capacity Plan Down

A capacity plan that lives in someone's head is re-derived under pressure. A one-page document per critical queue, reviewed before known peaks, keeps the numbers, the assumptions, and the decisions together — and makes it obvious when an assumption has changed.

## Capacity plan: orders queue (reviewed 2026-11-10)
Inputs (measured over last 30 days):
  normal peak λ = 60/s; S p50 = 0.8 s, p95 = 2.1 s; broker ceiling = 9,000 jobs/s (bench 2026-10)
Forecast: plateau 240/s 15:00-21:00; spike 600/s 18:00-18:15
Decisions:
  - pre-scale order pool to 35 pods at 14:30 (scheduled), KEDA max 50
  - recommendations moved to own pool, max 6 pods, allowed to lag
  - accept spike backlog ≤ 250k; expected drain ≈ 34 min
Dependencies: orders DB max_connections 600 (need ≤ 400); payment API 500 req/s (need 240)
Alerts: queue wait p95 > 5 min for 10 min -> page; drain ETA > 60 min -> page
Assumptions that invalidate this plan: S p50 > 1.0 s in load test; forecast > 300/s plateau

The last line is the most useful: it names what would make the plan wrong, so the team knows exactly which numbers to re-check before the day. After the event, compare the forecast with what actually happened — peak λ, measured S at peak, backlog size, drain time — and fold the differences into the next plan. Two or three cycles of this turn capacity planning from guesswork into a routine with known error bars.

Cost: Planning for the Trough as Well as the Peak

Capacity planning is usually framed around the peak, but most of the bill is spent in the trough. A fleet sized for Black Friday that runs at that size all year wastes most of its cost; a fleet that scales with load pays roughly for the work it does. Three levers reduce cost without touching the peak plan:

  • Scale to a floor, not to peak. Keep enough pods for the normal daily peak at 70% utilisation and let autoscaling add the rest; for queues with long idle periods, scaling workers to zero removes the floor entirely.
  • Use interruptible capacity for interruptible work. Jobs that are idempotent and short tolerate preemption well; running them on spot or preemptible instances costs a fraction of on-demand, as described in cutting worker costs with spot instances.
  • Improve service time. A 20% reduction in S is a 20% reduction in required slots at every load level, peak and trough alike — often the cheapest capacity available.

FAQ

How much headroom is enough? Plan forecast peak at 60–75% utilisation. Below 60% you pay for idle capacity; above 75–80%, normal burstiness produces visible queue waits.

Should I size on average or p99 service time? Throughput capacity uses mean service time (Little's Law is about averages). Latency targets need the distribution: high variance in S raises queue wait at the same utilisation, so check waits in a load test.

How often should capacity plans be revisited? Before every known peak, after any change that alters service time significantly (a new dependency, a heavier job type), and quarterly against growth.

How do I forecast λ when the business only gives me revenue or user targets? Convert through ratios you can measure: jobs per order, orders per active user, emails per signup. Pull thirty days of history for each ratio, check it is stable across normal peaks, and multiply the business forecast through it. Keep the ratio in the capacity plan; when a product change alters it (a new confirmation email per order, say), the plan shows exactly which number moved.

What if service time is highly variable? Mixed workloads — most jobs take 100 ms, a few take a minute — make averages misleading and queue waits unpredictable, because long jobs occupy slots that short jobs are waiting for. Separate them onto different queues with their own pools, and size each with its own λ and S. The split usually improves both throughput and tail latency more than adding workers would.

Can a queue be too big? Yes — an unbounded backlog hides a permanent shortfall until the data is useless (a confirmation email ten hours late). Set a maximum acceptable wait and alert well before it.

Related