Scaling Workers to Zero

Many queues are idle most of the day: nightly exports, monthly billing runs, a webhook queue for a feature few customers use. Keeping workers running for them costs money for nothing. This guide scales worker deployments down to zero when their queue is empty and back up when work arrives, as part of Horizontal Worker Scaling in Backend Frameworks & Worker Scaling.

Problem Statement

A platform runs 22 worker deployments on Kubernetes, each with at least two replicas for availability — 44 pods, many requesting 1–2 GB of memory. Usage data shows 15 of those queues are empty more than 90% of the time; their workers idle around the clock. The finance team wants the idle cost gone. The engineering concern is latency: some of these queues handle jobs a user waits for (report generation), and a job arriving when no worker exists must still start within an acceptable time. You want idle queues to cost nothing, scale-up fast enough for their latency needs, no lost scheduled jobs, and a clear rule for which queues should not scale to zero.

Prerequisites

  • Kubernetes with KEDA 2.x installed (or an equivalent event-driven autoscaler).
  • A queue KEDA can read: RabbitMQ, Redis lists or streams, SQS, Kafka, Postgres (via a SQL scaler), or a Prometheus metric.
  • Measured pod startup time for each worker image (image pull plus application boot).
  • The latency objective for each queue.

Step 1 — Decide Which Queues Qualify

Scale-to-zero trades idle cost for cold-start latency: the first job after an idle period waits for a pod to be scheduled, pulled, and started. Classify queues before configuring anything.

Queue Idle share Latency objective Cold start Scale to zero?
nightly-exports 95% 30 min ~40 s Yes
monthly-billing 99% hours ~40 s Yes
report-generation 85% 60 s ~40 s Yes, with a warm path (Step 5)
password-reset-email 20% 5 s ~40 s No
payments 30% 10 s ~40 s No

Queues with objectives shorter than cold start, or with frequent small bursts, keep a minimum of one or two replicas.

Idle share versus latency objective Queues that are idle most of the time and whose latency objective is longer than the roughly forty-second cold start, such as nightly exports and monthly billing, scale to zero. Queues with objectives shorter than cold start, such as password reset emails and payments, keep a warm minimum of replicas regardless of how idle they are. Which queues can scale to zero objective > cold start nightly exports, billing, report generation minReplicaCount: 0 objective < cold start password reset emails, payments minReplicaCount: 1-2 Measure cold start per image; it is often longer than people assume.

Step 2 — Configure KEDA with an Activation Threshold

KEDA separates two decisions: activation (0 ↔ 1 replica) and scaling (1 ↔ N, via an HPA it manages). activationValue sets how much work wakes the deployment; value sets the target per replica once running.

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata: { name: nightly-exports }
spec:
  scaleTargetRef: { name: exports-worker }
  minReplicaCount: 0
  maxReplicaCount: 10
  pollingInterval: 15          # check the queue every 15 s
  cooldownPeriod: 300          # wait 5 min of emptiness before scaling to 0
  triggers:
    - type: rabbitmq
      metadata:
        queueName: exports
        mode: QueueLength
        value: "20"            # target 20 messages per replica when active
        activationValue: "0"   # any message wakes the deployment
      authenticationRef: { name: rabbitmq-auth }

pollingInterval bounds how quickly KEDA notices the first job; total pickup latency after idle is roughly pollingInterval + pod start time. The cooldown prevents flapping between 0 and 1 when jobs arrive every few minutes. Scaling on queue length generally is covered in scaling workers with KEDA on queue length.

Step 3 — Count In-Flight Work, Not Just Waiting Messages

If the scaler counts only ready messages, a single long-running job can be in progress while the queue shows zero waiting — and KEDA scales the deployment to zero, killing the job. Make the metric include unacknowledged (in-flight) messages, or add a second trigger for them.

triggers:
  - type: prometheus
    metadata:
      serverAddress: http://prometheus.monitoring:9090
      query: |
        sum(rabbitmq_queue_messages_ready{queue="exports"})
        + sum(rabbitmq_queue_messages_unacked{queue="exports"})
      threshold: "20"
      activationThreshold: "0"

For Redis-based frameworks, count the active set too (BullMQ active, Sidekiq busy jobs, Celery's reserved/active tasks via an exporter). The cooldown also helps, but relying on it alone is how long jobs get killed at minute five of a ten-minute run.

Count work in progress A single export job is picked up, so the ready count drops to zero while the job runs for ten minutes. A scaler that watches only ready messages sees zero, waits out its five-minute cooldown, and scales to zero, killing the job at minute five. A scaler that counts ready plus unacknowledged messages sees one until the job finishes and scales to zero only afterwards. One 10-minute export job, cooldown 5 minutes ready only job running, metric = 0 scaled to 0 at 5 min: job killed ready + unacked job running, metric = 1 then 0 Scale-to-zero must only happen when there is no work waiting and none in progress.

Step 4 — Keep Scheduled and Delayed Jobs from Being Stranded

Scheduled jobs live outside the ready queue until due: Celery ETA tasks held by workers, Sidekiq's scheduled set, BullMQ's delayed set. With no workers running, some frameworks stop moving due jobs into the queue at all.

  • Celery holds eta/countdown tasks inside workers (reserved but unacked). With zero workers, they sit in the broker until a worker starts — and the scaler may not count them. Prefer a scheduler (beat or a database table) that enqueues at the due time instead of long countdowns.
  • Sidekiq moves scheduled jobs using a poller running inside Sidekiq processes. Zero processes means no poller; scheduled jobs wait until something wakes a worker.
  • BullMQ promotes delayed jobs through workers' Redis calls; with no workers, delayed jobs wait.

The fix is to include "due soon" work in the scaling metric, or keep a tiny always-on scheduler process that enqueues due jobs into the ready queue so the scaler sees them.

# BullMQ: wake when delayed jobs are due within the next minute (metric from an exporter)
query: |
  sum(bullmq_queue_jobs{queue="exports", state=~"waiting|active|prioritized"})
  + sum(bullmq_delayed_due_within_60s{queue="exports"})

Step 5 — Hide Cold Starts for User-Facing Queues

For report generation, a user clicking "export" should not wait 40 seconds for a pod. Two options keep scale-to-zero without the latency:

  1. Pre-warm on intent: when a user opens the report page, the API sets a short-lived "warm" signal that the scaler reads as one unit of demand, so a worker is starting before the user clicks.
  2. Faster cold start: smaller images, pre-pulled images on nodes (a DaemonSet that pulls the worker image), and lighter startup (lazy imports) can cut cold start from 40 seconds to under 10.
# API: record intent; the scaler query adds count(warm keys) to demand
redis.set(f"warm:reports:{user_id}", 1, ex=120)
Where cold-start time goes After a job arrives at an idle queue, KEDA notices it within the 15-second polling interval, the pod is scheduled in a couple of seconds, the image pull takes about fifteen seconds, and the application boots in about eight, for roughly forty seconds before the job starts. Pre-pulled images remove the pull, lazy imports shorten boot, and pre-warming on user intent starts the pod before the job exists. First job after idle: time to start default KEDA poll ≤ 15 s image pull ~15 s boot ~8 s optimised pre-warmed on intent, image pre-pulled, lazy imports Measure each segment; the image pull is usually the easiest to remove.

Measure the result: p95 time from enqueue to start for the first job after idle periods, compared with the queue's objective. Queue-wait measurement is covered in measuring queue wait time with enqueue timestamps.

Step 6 — Check the Savings and the Side Effects

After rollout, compare pod-hours and cost per queue, and check that latency objectives still hold.

# Replica-hours per worker deployment over a week
sum by (deployment) (avg_over_time(kube_deployment_status_replicas{deployment=~".*-worker"}[7d])) * 24 * 7

# First-job-after-idle wait (jobs whose wait exceeded 20 s, as a share)
sum(rate(job_queue_wait_seconds_bucket{le="+Inf",queue="exports"}[1d]))
  - sum(rate(job_queue_wait_seconds_bucket{le="20",queue="exports"}[1d]))

In the scenario, 15 deployments at zero most of the day cut worker pod-hours by about 60%, with report generation's p95 wait unchanged after pre-warming.

Verification

# Queue empty: deployment should reach 0 after the cooldown
kubectl get scaledobject nightly-exports -o jsonpath='{.status.conditions[?(@.type=="Active")].status}'
kubectl get deploy exports-worker -o jsonpath='{.status.replicas}'      # expect 0

# Enqueue one job: a pod should start within pollingInterval + startup
python enqueue_export.py && time kubectl wait --for=condition=available deploy/exports-worker --timeout=120s

Then run a single 10-minute job and confirm the deployment does not scale to zero until it finishes.

Gotchas & Edge Cases

Cluster autoscaler scale-up. If no node has room, a new node must start too, adding minutes to cold start. Keep headroom or use a warm node pool for latency-sensitive queues.

Connection storms on wake. Many deployments waking together (a batch that touches several queues) can hit the database at once. Stagger or cap maxReplicaCount.

Monitoring gaps. Workers at zero emit no metrics; alerts that expect worker metrics (heartbeat missing) must be scoped to active deployments.

Secrets and config mounted at startup. A worker that fetches secrets from a vault or reads a large config map on boot adds that latency to every wake. Cache what you can in the image or use a sidecar that is already warm on the node.

Readiness vs availability. A pod that is scheduled but still booting is not processing; measure pickup latency, not replica count.

FAQ

Can HPA alone scale to zero? The standard HPA does not scale to zero (without an alpha feature gate); KEDA handles the 0 ↔ 1 step and hands 1 ↔ N to an HPA it manages.

What about serverless workers instead? Lambda or Cloud Run consume queues with scale-to-zero built in and very fast cold starts for small functions — see processing SQS with AWS Lambda. For heavier workers with large images, KEDA on Kubernetes is often simpler.

Does scale-to-zero affect graceful shutdown? Scaling to zero removes the last replica with the same SIGTERM and grace period as any scale-in, so the drain configuration from graceful shutdown & worker deployments applies unchanged. Counting in-flight work (Step 3) is what ensures the last pod is not removed while it still has a job.

What if the scaler itself fails? If KEDA cannot read the metric, it keeps the current replica count by default (and can use a configured fallback). Alert on scaler errors; a broken scaler at zero replicas means jobs accumulate with nobody processing them.

How long should the cooldown be? Longer than the typical gap between jobs during active periods, so the deployment does not bounce between 0 and 1; 5–15 minutes is common.

Related