Scaling RQ Workers in Production

RQ (Redis Queue) is deliberately simple: one worker process takes one job at a time, forks a child to run it, and records the result in Redis. That simplicity makes it easy to start with and easy to reason about, but it also means scaling RQ is mostly about running the right number of processes in the right places. This guide covers how, as part of RQ vs Celery for Python in Backend Frameworks & Worker Scaling.

Problem Statement

A Django application uses RQ for emails, report exports, and nightly data imports. Four worker processes run under systemd on one VM, all listening to high,default,low. During month-end, report exports take 20 minutes each and occupy every worker, so password-reset emails wait behind them. Some exports are killed by the default 180-second timeout and retried forever by a cron script. After a VM reboot, jobs that were running show as "started" in the dashboard indefinitely. The team wants predictable latency for urgent jobs, a way to run more workers across machines, and clean recovery of interrupted jobs โ€” without switching frameworks.

Prerequisites

  • RQ 1.16 or 2.x with Redis 6 or newer, configured with maxmemory-policy noeviction.
  • A process manager: systemd, supervisord, or Kubernetes.
  • Job durations per job type (from job.started_at and job.ended_at, or your metrics).
  • The ability to set per-job job_timeout and result_ttl at enqueue time.

Step 1 โ€” Understand RQ's Execution Model

Each rq worker process handles exactly one job at a time. For each job, the default Worker class forks a "work horse" child, runs the job in it, and waits. The fork gives isolation โ€” a job that leaks memory or crashes the interpreter only kills the child โ€” at the cost of a few milliseconds and copy-on-write memory per job.

How an RQ worker runs a job A worker process blocks on Redis until a job is available, then forks a work horse child process that executes the job while the parent monitors it and sends heartbeats. When the child exits, the parent records success or failure in Redis and fetches the next job. One worker runs one job at a time, so total concurrency equals the number of worker processes. One worker, one job, one fork Redis high ยท default ยท low BLPOP worker process heartbeat, timeout watch fork work horse runs the job, exits result or failure written back, then next job Concurrency = number of worker processes

Consequences for scaling: concurrency is the number of processes; memory is roughly (parent + peak child) ร— processes; and a CPU core can host several workers only if the jobs mostly wait on I/O. For I/O-heavy workloads where fork overhead dominates, SimpleWorker runs jobs in-process without forking โ€” faster, but a crashing job takes the worker with it, and memory leaks accumulate.

Step 2 โ€” Run Many Workers with worker-pool or a Process Manager

RQ 1.14 added rq worker-pool, which starts and supervises N workers from one command:

rq worker-pool high default low --num-workers 8 --url redis://queue-redis:6379/0

Under systemd, a template unit gives the same result with per-instance restarts and logs:

# /etc/systemd/system/rq-worker@.service
[Service]
ExecStart=/srv/app/.venv/bin/rq worker --with-scheduler high default low --name %H-%i
Restart=always
KillSignal=SIGTERM
TimeoutStopSec=600

systemctl enable --now rq-worker@{1..8} starts eight. Give each worker a unique, stable name (host-index) so the dashboard and the registries are readable. In Kubernetes, run one worker per container and scale replicas, which lets the cluster place workers across nodes and restart them independently. Whatever the mechanism, spread workers over at least two hosts so one failure does not stop processing.

Step 3 โ€” Separate Urgent Work from Slow Work

RQ workers check their queues in the order given: rq worker high default low always takes from high first. That provides priority but not isolation โ€” if every worker is busy with a 20-minute export from low, a new high job still waits. Fix it by dedicating workers:

Shared vs dedicated RQ workers On the left, four workers listen to high, default and low; all four are running long exports, so an urgent email job waits until one finishes. On the right, two workers listen only to high and are idle and ready, while four workers listen to default and low and handle the exports. Urgent latency no longer depends on how many slow jobs are running. Priority order is not isolation shared: high default low all four running 20-minute exports email job waits dedicated pools high only: free default low: exports email runs at once Size the high pool for its peak, not the average.
rq worker-pool high --num-workers 2          # urgent only
rq worker-pool high default low --num-workers 4   # everything else, high first

Letting the general pool also listen to high means urgent work gets extra capacity when it is available, while the dedicated pool guarantees a floor. Put very long jobs (imports, exports) on their own queue with their own workers so they never occupy the general pool at all.

Step 4 โ€” Set Timeouts, Result TTLs, and Retries per Job

RQ's default job_timeout is 180 seconds; a longer job is killed and moved to the failed registry. Set it per job type, from measured p99 duration plus a margin:

from rq import Retry

queue.enqueue(export_report, report_id,
              job_timeout="45m",
              result_ttl=3600,          # keep successful results one hour
              failure_ttl=7 * 86400,    # keep failures a week for inspection
              retry=Retry(max=3, interval=[30, 120, 600]))

Keep result_ttl short or zero for jobs whose return value nobody reads โ€” results live in Redis and add up quickly at high volume. Use Retry for transient failures instead of a cron script re-enqueueing failed jobs, which cannot distinguish a timeout from a bug. For the patterns behind choosing intervals, see retry strategies & backoff.

Step 5 โ€” Recover Abandoned Jobs After Crashes

When a worker process is killed without a chance to clean up (OOM, VM reboot, kill -9), its current job stays in the StartedJobRegistry. RQ's maintenance task โ€” run periodically by any worker โ€” moves jobs whose worker heartbeat has expired to the FailedJobRegistry as "abandoned". They are not retried automatically unless you act:

from rq.registry import FailedJobRegistry

registry = FailedJobRegistry(queue=queue)
for job_id in registry.get_job_ids():
    job = queue.fetch_job(job_id)
    if job and "abandoned" in (job.exc_info or "").lower():
        registry.requeue(job_id)
Where an interrupted RQ job ends up A job starts in the queue, moves to the started registry when a worker takes it, and then ends in the finished registry on success or the failed registry on an exception or timeout. If the worker dies, the job stays in started until the heartbeat expires, when maintenance moves it to failed marked as abandoned. A scheduled sweep requeues abandoned jobs of types that are safe to run twice. Job registries and the abandoned-job path queue started finished (result_ttl) failed (failure_ttl) error, timeout, or abandoned sweep requeues abandoned, idempotent jobs

Run a sweep like this on a schedule, and only for job types that are safe to run twice. Handle SIGTERM properly too: RQ's warm shutdown finishes the current job before exiting, so give the process manager a stop timeout longer than your longest job โ€” the details are in draining RQ workers safely.

Step 6 โ€” Autoscale on Queue Length

Because one worker equals one concurrent job, the number of workers you need follows directly from arrival rate and job duration: workers โ‰ˆ arrival_rate ร— avg_duration / target_utilisation. For bursty queues, scale on backlog instead. KEDA's Redis list scaler reads the queue length directly (RQ queues are the Redis lists rq:queue:<name>):

triggers:
  - type: redis
    metadata:
      address: queue-redis:6379
      listName: rq:queue:default
      listLength: "20"      # target waiting jobs per worker replica

Scale each dedicated pool on its own queue. Remember that queue length excludes running jobs, so a scaler that sees an empty list may remove a worker that is mid-job; rely on graceful shutdown and a long termination grace period. The general approach is covered in scaling workers with KEDA on queue length.

Step 7 โ€” Watch the Redis Side

Every worker holds a connection, sends heartbeats, and updates registries. At a few hundred workers, Redis CPU and connection count become visible. Keep result_ttl low, avoid storing large return values, cap the number of workers per Redis instance at what load tests show it handles, and move very high-volume queues to their own Redis instance. If you need thousands of concurrent I/O-bound jobs, RQ's process-per-job model is the wrong fit; that is the point to look at Celery with gevent, Dramatiq with threads, or an async framework โ€” see Dramatiq vs Celery.

Verification

  • rq info --interval 5 shows the expected number of workers per queue, spread across hosts.
  • Enqueue a high job while all general workers run exports: it starts within a second.
  • A job exceeding its job_timeout lands in the failed registry with a timeout error, and its retry schedule matches the configuration.
  • After kill -9 of a worker, the job appears as abandoned within the heartbeat window and is requeued by the sweep.
  • Redis memory from results stays flat over a day of normal traffic.

Gotchas & Edge Cases

Fork and database connections. A work horse inherits the parent's open sockets. Close database connections before forking (Django: django.db.connections.close_all() in a custom worker class) or each child can corrupt the parent's connection state.

macOS and fork safety. Forking after Objective-C libraries initialise can crash on macOS; set OBJC_DISABLE_INITIALIZE_FORK_SAFETY=YES in development or use SimpleWorker locally.

Scheduler duplication. --with-scheduler lets any worker move scheduled jobs to their queues, coordinated by a lock. It is safe to enable on many workers, but enqueueing periodic jobs needs a single source such as rq-scheduler or a cron job with a lock.

Job arguments are pickled. RQ serialises with pickle by default. Pass IDs rather than model instances, and never enqueue from untrusted input without a JSON serializer.

FAQ

How many RQ workers can one machine run? As many as memory allows for I/O-bound jobs โ€” often 10โ€“30 per core โ€” and about one per core for CPU-bound jobs. Measure memory per worker at peak job size and leave headroom.

Should I use SimpleWorker to avoid fork overhead? Only for short, trusted, well-behaved jobs where the fork cost dominates. You lose crash isolation and per-job memory cleanup, so restart workers periodically with --max-jobs.

When should I move from RQ to Celery? When you need workflows, high per-process concurrency, brokers other than Redis, or fine-grained routing that RQ's queue lists cannot express. The runbook is in migrating from RQ to Celery.

Related