Calculating Worker Count with Little's Law
"How many workers do we need?" has a precise answer once you measure two numbers, and this guide derives it step by step as part of Capacity Planning for Job Queues in Observability & Monitoring for Job Queues. Little's Law gives the average number of busy worker slots; a utilisation target turns that into a fleet size; and a standard queueing estimate tells you what queue wait to expect at that size.
Problem Statement
A document-processing service runs 12 worker pods with 4 concurrent slots each. At its daily peak it receives 30 documents per second, and queue wait climbs from under a second to several minutes for about an hour every afternoon. The team's instinct is to double the pods, but the finance review wants a justified number, and the platform team suspects that concurrency per pod, not pod count, is the real constraint. You need to compute the required slots from measured data, choose a target utilisation that keeps queue wait under 10 seconds at p95, and verify the result against production.
Prerequisites
- Metrics for enqueue rate and job duration per queue (histograms for duration), as in Prometheus Metrics for Workers.
- A queue-wait metric, ideally measured from an enqueue timestamp on each job.
- At least a week of history covering normal peaks.
- A clear definition of "slot" for your framework: prefork processes, threads, goroutines, or async concurrency — whatever limits how many jobs one pod runs at once.
Step 1 — Measure Arrival Rate at Peak
Use the rate of enqueues (not completions — completions are capped by capacity and understate demand during a backlog). Take a high percentile of the 5-minute rate over the peak hour across several days.
# λ: 95th percentile of the 5-minute enqueue rate during weekday peaks, last 7 days
quantile_over_time(0.95,
sum(rate(jobs_enqueued_total{queue="documents"}[5m]))[7d:5m]
)
# -> 30.4 jobs/s
If you only have completion counts, measure λ on a day with no backlog, when completions equal arrivals.
Step 2 — Measure Service Time Correctly
Service time is how long a slot is occupied per job: from the moment a worker takes the job until it is free for the next one. That includes the handler, but also acknowledgement, result storage, and any per-job overhead in the worker loop. The job-duration histogram usually captures the handler only; if the gap matters, compare busy slot-seconds / completed jobs over a window.
# Mean service time (Little's Law uses the mean)
sum(rate(job_duration_seconds_sum{queue="documents"}[1h]))
/ sum(rate(job_duration_seconds_count{queue="documents"}[1h]))
# -> 1.9 s
# Variability: p95 / mean (used in Step 5)
histogram_quantile(0.95, sum by (le) (rate(job_duration_seconds_bucket{queue="documents"}[1h])))
# -> 5.2 s
Measure S during the peak, not overnight: under load, contention in shared dependencies usually makes jobs slower, and that slower S is the one that matters for peak sizing.
Step 3 — Compute Busy Slots with Little's Law
Little's Law states that the average number of items in a system equals the arrival rate times the average time each spends in it: L = λ × W. Applied to the service part of a worker fleet, the average number of busy slots is λ × S.
lam = 30.4 # jobs/s at peak
S = 1.9 # mean seconds per job
busy = lam * S # 57.8 slots busy on average
current = 12 * 4 # 48 slots
print(busy, current, busy / current) # 57.8 48 1.20 -> demand exceeds capacity by 20%
This already explains the afternoon backlog: at peak, the fleet needs about 58 slots busy on average but has 48, so the queue grows at λ - current/S ≈ 30.4 - 25.3 = 5.1 jobs per second until demand drops. Doubling pods would fix it — with a lot of spare capacity — but Little's Law gives the minimum, and the next step picks a sensible margin above it.
Step 4 — Apply a Utilisation Target
Running at 100% utilisation means any burst above the average creates a backlog that cannot drain until load falls. Divide busy slots by a target utilisation to leave room for variance.
import math
def slots_for(lam: float, S: float, target_util: float) -> int:
return math.ceil(lam * S / target_util)
for u in (0.9, 0.8, 0.7, 0.6):
s = slots_for(30.4, 1.9, u)
print(f"target {u:.0%}: {s} slots = {math.ceil(s / 4)} pods at 4 slots")
# target 90%: 65 slots = 17 pods
# target 80%: 73 slots = 19 pods
# target 70%: 83 slots = 21 pods
# target 60%: 97 slots = 25 pods
Which target to pick depends on how variable arrivals and service times are, and on the queue-wait objective. Step 5 turns that into a number rather than a guess.
Step 5 — Estimate Queue Wait with an M/M/c Model
The M/M/c model (Poisson arrivals, exponential service times, c servers) gives the probability that an arriving job must wait and the expected wait. Real workloads are rarely exactly exponential, but the model captures the key non-linearity and is a good first estimate; the Erlang C formula computes it.
def erlang_c(c: int, a: float) -> float:
"""Probability an arrival waits. a = λ·S (offered load in slots), c = slots."""
if a >= c:
return 1.0
s, term = 0.0, 1.0
for k in range(c):
if k > 0:
term *= a / k
s += term
term *= a / c # a^c / c!
top = term * c / (c - a)
return top / (s + top)
def mean_wait(c: int, lam: float, S: float) -> float:
a = lam * S
return erlang_c(c, a) * S / (c - a)
for c in (65, 73, 83, 97):
print(c, f"P(wait)={erlang_c(c, 30.4 * 1.9):.2f}", f"Wq={mean_wait(c, 30.4, 1.9):.2f}s")
# 65 P(wait)=0.40 Wq=0.10s
# 73 P(wait)=0.11 Wq=0.01s
# 83 ...
The model predicts tiny mean waits at 65–73 slots, yet production showed minutes of wait with bursts. The gap is variance: arrivals come in bursts (batch uploads) and service times are heavy-tailed (p95 is 2.7× the mean). A practical correction (the Allen–Cunneen approximation) multiplies the M/M/c wait by (Ca² + Cs²) / 2, where Ca and Cs are the coefficients of variation of inter-arrival and service times. With bursty arrivals and heavy-tailed service times this factor is often 3–10, which is why real fleets need a 70% target where the pure model suggests 90% would do.
Step 6 — Decide Between More Pods and More Slots per Pod
83 slots can be 21 pods × 4 or 11 pods × 8. Which is right depends on what the job is bound by. For I/O-bound jobs (waiting on storage, APIs), more slots per pod is cheaper — idle waiting costs little CPU. For CPU-bound jobs, slots beyond the pod's cores just time-slice and increase S. Measure CPU per busy slot:
# CPU seconds per job: if ~equal to S, the job is CPU-bound; if much lower, I/O-bound
sum(rate(container_cpu_usage_seconds_total{pod=~"doc-worker-.*"}[10m]))
/ sum(rate(job_duration_seconds_count{queue="documents"}[10m]))
# -> 0.35 CPU-seconds per job vs S = 1.9 s -> mostly I/O-bound
At 0.35 CPU-seconds per job, each busy slot uses about 18% of a core, so a 2-core pod can run 8 slots comfortably. That halves the pod count to 11 and the platform team's suspicion is confirmed.
The general method is in right-sizing worker concurrency per CPU.
Verification
Deploy the new size before a peak and compare prediction with measurement:
# Predicted busy slots (λ·S) vs observed busy slots
sum(rate(jobs_enqueued_total{queue="documents"}[5m])) * 1.9
sum(worker_busy_slots{queue="documents"})
# Observed utilisation and the objective
sum(worker_busy_slots{queue="documents"}) / sum(worker_total_slots{queue="documents"})
histogram_quantile(0.95, sum by (le) (rate(job_queue_wait_seconds_bucket{queue="documents"}[5m])))
Predicted and observed busy slots should agree within about 10%; if observed is higher, service time under the new configuration grew (check CPU throttling). Queue wait p95 should stay under the 10-second objective through the peak.
Gotchas & Edge Cases
Using completion rate as λ during a backlog. Completions are capped by capacity, so the calculation returns exactly the capacity you already have. Use enqueues.
Ignoring retries. Retried jobs arrive again; include them in λ (enqueue counters that count retries do this automatically).
Mixed job types on one queue. A queue with 100 ms and 60 s jobs has a meaningless average. Split into queues by job type and size each, or at least compute λ × S per type and sum.
Service time that depends on concurrency. If S grows as you add slots (shared database contention), Little's Law still holds but you must measure S at the target concurrency. Load-test it — see load testing queue throughput.
FAQ
Is Little's Law only for steady state? It holds for long-run averages in any stable system, regardless of distributions. It does not describe a transient backlog — use drain-time arithmetic for that, as in forecasting backlog drain time.
Why not just autoscale and skip the maths? Autoscaling needs a min, a max, and a target, and all three come from this calculation. It also reacts in minutes, so the floor must cover predictable peaks.
Can Little's Law estimate queue wait directly?
Yes, if you know queue length: W_q = L_q / λ. With 150 jobs waiting and 30 per second arriving, the average wait is 5 seconds — a quick sanity check during incidents.
Related
- Capacity Planning for Job Queues — the wider planning process.
- Load Testing Queue Throughput — measure S and the ceiling at target load.
- Horizontal Worker Scaling — turning the number into an autoscaling policy.
- Measuring Queue Wait Time with Enqueue Timestamps — the wait metric this guide validates against.