Scaling Workers to Zero
Many queues are idle most of the day: nightly exports, monthly billing runs, a webhook queue for a feature few customers use. Keeping workers running for them costs money for nothing. This guide scales worker deployments down to zero when their queue is empty and back up when work arrives, as part of Horizontal Worker Scaling in Backend Frameworks & Worker Scaling.
Problem Statement
A platform runs 22 worker deployments on Kubernetes, each with at least two replicas for availability — 44 pods, many requesting 1–2 GB of memory. Usage data shows 15 of those queues are empty more than 90% of the time; their workers idle around the clock. The finance team wants the idle cost gone. The engineering concern is latency: some of these queues handle jobs a user waits for (report generation), and a job arriving when no worker exists must still start within an acceptable time. You want idle queues to cost nothing, scale-up fast enough for their latency needs, no lost scheduled jobs, and a clear rule for which queues should not scale to zero.
Prerequisites
- Kubernetes with KEDA 2.x installed (or an equivalent event-driven autoscaler).
- A queue KEDA can read: RabbitMQ, Redis lists or streams, SQS, Kafka, Postgres (via a SQL scaler), or a Prometheus metric.
- Measured pod startup time for each worker image (image pull plus application boot).
- The latency objective for each queue.
Step 1 — Decide Which Queues Qualify
Scale-to-zero trades idle cost for cold-start latency: the first job after an idle period waits for a pod to be scheduled, pulled, and started. Classify queues before configuring anything.
| Queue | Idle share | Latency objective | Cold start | Scale to zero? |
|---|---|---|---|---|
| nightly-exports | 95% | 30 min | ~40 s | Yes |
| monthly-billing | 99% | hours | ~40 s | Yes |
| report-generation | 85% | 60 s | ~40 s | Yes, with a warm path (Step 5) |
| password-reset-email | 20% | 5 s | ~40 s | No |
| payments | 30% | 10 s | ~40 s | No |
Queues with objectives shorter than cold start, or with frequent small bursts, keep a minimum of one or two replicas.
Step 2 — Configure KEDA with an Activation Threshold
KEDA separates two decisions: activation (0 ↔ 1 replica) and scaling (1 ↔ N, via an HPA it manages). activationValue sets how much work wakes the deployment; value sets the target per replica once running.
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata: { name: nightly-exports }
spec:
scaleTargetRef: { name: exports-worker }
minReplicaCount: 0
maxReplicaCount: 10
pollingInterval: 15 # check the queue every 15 s
cooldownPeriod: 300 # wait 5 min of emptiness before scaling to 0
triggers:
- type: rabbitmq
metadata:
queueName: exports
mode: QueueLength
value: "20" # target 20 messages per replica when active
activationValue: "0" # any message wakes the deployment
authenticationRef: { name: rabbitmq-auth }
pollingInterval bounds how quickly KEDA notices the first job; total pickup latency after idle is roughly pollingInterval + pod start time. The cooldown prevents flapping between 0 and 1 when jobs arrive every few minutes. Scaling on queue length generally is covered in scaling workers with KEDA on queue length.
Step 3 — Count In-Flight Work, Not Just Waiting Messages
If the scaler counts only ready messages, a single long-running job can be in progress while the queue shows zero waiting — and KEDA scales the deployment to zero, killing the job. Make the metric include unacknowledged (in-flight) messages, or add a second trigger for them.
triggers:
- type: prometheus
metadata:
serverAddress: http://prometheus.monitoring:9090
query: |
sum(rabbitmq_queue_messages_ready{queue="exports"})
+ sum(rabbitmq_queue_messages_unacked{queue="exports"})
threshold: "20"
activationThreshold: "0"
For Redis-based frameworks, count the active set too (BullMQ active, Sidekiq busy jobs, Celery's reserved/active tasks via an exporter). The cooldown also helps, but relying on it alone is how long jobs get killed at minute five of a ten-minute run.
Step 4 — Keep Scheduled and Delayed Jobs from Being Stranded
Scheduled jobs live outside the ready queue until due: Celery ETA tasks held by workers, Sidekiq's scheduled set, BullMQ's delayed set. With no workers running, some frameworks stop moving due jobs into the queue at all.
- Celery holds
eta/countdowntasks inside workers (reserved but unacked). With zero workers, they sit in the broker until a worker starts — and the scaler may not count them. Prefer a scheduler (beat or a database table) that enqueues at the due time instead of long countdowns. - Sidekiq moves scheduled jobs using a poller running inside Sidekiq processes. Zero processes means no poller; scheduled jobs wait until something wakes a worker.
- BullMQ promotes delayed jobs through workers' Redis calls; with no workers, delayed jobs wait.
The fix is to include "due soon" work in the scaling metric, or keep a tiny always-on scheduler process that enqueues due jobs into the ready queue so the scaler sees them.
# BullMQ: wake when delayed jobs are due within the next minute (metric from an exporter)
query: |
sum(bullmq_queue_jobs{queue="exports", state=~"waiting|active|prioritized"})
+ sum(bullmq_delayed_due_within_60s{queue="exports"})
Step 5 — Hide Cold Starts for User-Facing Queues
For report generation, a user clicking "export" should not wait 40 seconds for a pod. Two options keep scale-to-zero without the latency:
- Pre-warm on intent: when a user opens the report page, the API sets a short-lived "warm" signal that the scaler reads as one unit of demand, so a worker is starting before the user clicks.
- Faster cold start: smaller images, pre-pulled images on nodes (a DaemonSet that pulls the worker image), and lighter startup (lazy imports) can cut cold start from 40 seconds to under 10.
# API: record intent; the scaler query adds count(warm keys) to demand
redis.set(f"warm:reports:{user_id}", 1, ex=120)
Measure the result: p95 time from enqueue to start for the first job after idle periods, compared with the queue's objective. Queue-wait measurement is covered in measuring queue wait time with enqueue timestamps.
Step 6 — Check the Savings and the Side Effects
After rollout, compare pod-hours and cost per queue, and check that latency objectives still hold.
# Replica-hours per worker deployment over a week
sum by (deployment) (avg_over_time(kube_deployment_status_replicas{deployment=~".*-worker"}[7d])) * 24 * 7
# First-job-after-idle wait (jobs whose wait exceeded 20 s, as a share)
sum(rate(job_queue_wait_seconds_bucket{le="+Inf",queue="exports"}[1d]))
- sum(rate(job_queue_wait_seconds_bucket{le="20",queue="exports"}[1d]))
In the scenario, 15 deployments at zero most of the day cut worker pod-hours by about 60%, with report generation's p95 wait unchanged after pre-warming.
Verification
# Queue empty: deployment should reach 0 after the cooldown
kubectl get scaledobject nightly-exports -o jsonpath='{.status.conditions[?(@.type=="Active")].status}'
kubectl get deploy exports-worker -o jsonpath='{.status.replicas}' # expect 0
# Enqueue one job: a pod should start within pollingInterval + startup
python enqueue_export.py && time kubectl wait --for=condition=available deploy/exports-worker --timeout=120s
Then run a single 10-minute job and confirm the deployment does not scale to zero until it finishes.
Gotchas & Edge Cases
Cluster autoscaler scale-up. If no node has room, a new node must start too, adding minutes to cold start. Keep headroom or use a warm node pool for latency-sensitive queues.
Connection storms on wake. Many deployments waking together (a batch that touches several queues) can hit the database at once. Stagger or cap maxReplicaCount.
Monitoring gaps. Workers at zero emit no metrics; alerts that expect worker metrics (heartbeat missing) must be scoped to active deployments.
Secrets and config mounted at startup. A worker that fetches secrets from a vault or reads a large config map on boot adds that latency to every wake. Cache what you can in the image or use a sidecar that is already warm on the node.
Readiness vs availability. A pod that is scheduled but still booting is not processing; measure pickup latency, not replica count.
FAQ
Can HPA alone scale to zero? The standard HPA does not scale to zero (without an alpha feature gate); KEDA handles the 0 ↔ 1 step and hands 1 ↔ N to an HPA it manages.
What about serverless workers instead? Lambda or Cloud Run consume queues with scale-to-zero built in and very fast cold starts for small functions — see processing SQS with AWS Lambda. For heavier workers with large images, KEDA on Kubernetes is often simpler.
Does scale-to-zero affect graceful shutdown? Scaling to zero removes the last replica with the same SIGTERM and grace period as any scale-in, so the drain configuration from graceful shutdown & worker deployments applies unchanged. Counting in-flight work (Step 3) is what ensures the last pod is not removed while it still has a job.
What if the scaler itself fails? If KEDA cannot read the metric, it keeps the current replica count by default (and can use a configured fallback). Alert on scaler errors; a broken scaler at zero replicas means jobs accumulate with nobody processing them.
How long should the cooldown be? Longer than the typical gap between jobs during active periods, so the deployment does not bounce between 0 and 1; 5–15 minutes is common.
Related
- Horizontal Worker Scaling — scaling principles.
- Scaling Workers with KEDA on Queue Length — KEDA basics.
- Cutting Worker Costs with Spot Instances — the other cost lever.
- Capacity Planning for Job Queues — sizing the active range.