Migrating from Redis to a Postgres Job Queue

Moving a live job workload from Redis to Postgres is a data migration with traffic flowing through it, and this guide walks through doing it one job class at a time as part of Database-Backed Job Queues in Backend Frameworks & Worker Scaling. The approach works whether the source is Sidekiq, RQ, Celery-on-Redis, or BullMQ, and whether the target is Solid Queue, River, Oban, Procrastinate, or a hand-built table.

Problem Statement

A Rails application runs 60 Sidekiq job classes at a peak of about 300 jobs per second. The team wants to drop Redis: it is the only component without point-in-time recovery, two incidents last year lost scheduled jobs during a failover, and several jobs suffer from the dual-write race between the database commit and the enqueue. Solid Queue is the target. A "big bang" switch is off the table — a bad cutover would stall payments and emails. You need a migration that moves jobs incrementally, never runs the same job on both systems, never drops a scheduled or retrying job, and can be reversed per job class in minutes.

Prerequisites

  • Both backends installable side by side: in Rails, Active Job allows a per-class adapter; in Python, a thin enqueue wrapper; in Node, a producer abstraction.
  • Capacity in Postgres for the added write load (roughly three writes per job) — check with a load test before starting, as in load testing queue throughput.
  • Dashboards for both systems: enqueue rate, processing rate, failures, and oldest job age per queue.
  • A feature-flag system that can switch per job class without a deploy.

Step 1 — Inventory Every Job Class and Its Redis-Specific Features

Each job class depends on some Redis-backend features that may not map one-to-one. Build the inventory from code and from live data, not from memory.

# script/job_inventory.rb — what runs, how often, and what it relies on
require "sidekiq/api"
stats = Hash.new { |h, k| h[k] = { enqueued: 0, scheduled: 0, retry: 0 } }
Sidekiq::Queue.all.each { |q| q.each { |j| stats[j.klass][:enqueued] += 1 } }
Sidekiq::ScheduledSet.new.each { |j| stats[j.klass][:scheduled] += 1 }
Sidekiq::RetrySet.new.each { |j| stats[j.klass][:retry] += 1 }

ApplicationJob.descendants.chain(Sidekiq::Job.descendants rescue []).each do |klass|
  opts = klass.respond_to?(:get_sidekiq_options) ? klass.get_sidekiq_options : {}
  puts [klass.name, stats[klass.name].values, opts.slice("queue", "retry", "unique_for",
        "lock", "throttle").inspect].flatten.join("\t")
end

Flag every class that uses a Sidekiq-specific feature: uniqueness (sidekiq-unique-jobs), throttling (sidekiq-throttled), batches (Sidekiq Pro), rate limiting (Enterprise), or custom middleware. Each needs an equivalent on the target — Solid Queue's limits_concurrency replaces most uniqueness and throttling uses, as covered in running Rails Solid Queue in production. Classes without such dependencies go first.

Migrate in waves, simplest first Wave one contains low-volume job classes with no Redis-specific features, such as report emails. Wave two contains high-volume plain jobs. Wave three contains jobs that depend on uniqueness, throttling, or batches and need equivalents on the new backend. Payment-critical jobs go last, after the others have proven the setup. 60 job classes, four waves wave 1: 22 low volume no special features wave 2: 18 high volume plain retries wave 3: 14 uniqueness, throttling, batches to re-map wave 4: 6 payments, billing after others prove out Each wave runs for at least a week of normal traffic before the next starts.

Step 2 — Run Both Worker Fleets Side by Side

Deploy the new backend's workers alongside the existing ones before moving any traffic. With nothing enqueued to them they sit idle, but you verify they boot, connect, heartbeat, and appear on dashboards.

# k8s: add the new worker deployment; keep sidekiq untouched
apiVersion: apps/v1
kind: Deployment
metadata: { name: jobs-solid-queue }
spec:
  replicas: 2
  template:
    spec:
      containers:
        - name: jobs
          image: registry.internal/app:${GIT_SHA}
          command: ["bin/jobs"]
          env:
            - { name: DB_POOL, value: "12" }
          resources:
            requests: { cpu: "500m", memory: "768Mi" }

Both fleets run the same application image. The only difference is which process starts, so every job class exists in both and can run on either — which is what makes per-class switching and rollback possible.

Step 3 — Switch Enqueue Per Job Class Behind a Flag

Route each job class's new enqueues to one backend based on a flag. In Rails, override the adapter per class; in other stacks, put the decision in your enqueue helper.

# app/jobs/application_job.rb
class ApplicationJob < ActiveJob::Base
  def self.queue_adapter
    if Flags.enabled?(:"jobs_on_solid_queue.#{name}")
      ActiveJob::QueueAdapters::SolidQueueAdapter.new
    else
      ActiveJob::QueueAdapters::SidekiqAdapter.new
    end
  end
end
# Python equivalent: one enqueue helper, per-task routing flag
def enqueue(task_name: str, *args, **kwargs):
    if flags.enabled(f"jobs_on_pg.{task_name}"):
        return pg_queue.enqueue(task_name, args, kwargs)       # e.g. Procrastinate
    return rq_queue.enqueue(task_name, *args, **kwargs)        # existing RQ

The flag decides only where a job is enqueued. A job already in Redis — queued, scheduled, or in the retry set — stays in Redis and is executed by Sidekiq. That property is what prevents a job from running on both systems: every job instance lives in exactly one backend for its whole life.

New jobs move; existing jobs stay put An enqueue router checks a per-class flag. With the flag on, new SendReceiptJob instances are inserted into Postgres and run by Solid Queue workers. Jobs of the same class that were already queued, scheduled, or retrying in Redis remain there and are finished by Sidekiq, so no job instance ever exists in both systems. Flag flips enqueue, not execution enqueue router flag per job class Postgres: new jobs Redis: in-flight backlog Solid Queue workers Sidekiq workers dashed: flag off, or jobs enqueued before the flip

Step 4 — Ramp One Class and Watch Both Sides

Turn the flag on for one low-risk class — for a percentage of enqueues if your flag system supports it — and compare behaviour between backends over a normal traffic cycle.

-- Target side: throughput, failures, and queue time for the migrated class
SELECT date_trunc('minute', finished_at) AS minute,
       count(*)                                                AS finished,
       count(*) FILTER (WHERE id IN (SELECT job_id FROM solid_queue_failed_executions)) AS failed,
       percentile_cont(0.95) WITHIN GROUP (ORDER BY finished_at - scheduled_at)       AS p95_wait
FROM solid_queue_jobs
WHERE class_name = 'SendReceiptJob' AND finished_at > now() - interval '1 hour'
GROUP BY 1 ORDER BY 1;

On the Sidekiq side, the same class's processing rate should fall as its Redis backlog drains, and its enqueue rate should drop to zero once the flag is at 100%. The combined processed count across both systems should match the enqueue count from your application metrics — a gap means jobs are being lost or duplicated somewhere.

Step 5 — Handle Scheduled and Retrying Jobs Explicitly

Jobs in Sidekiq's scheduled set and retry set can sit there for days. Leaving them to drain naturally is the safest option: they run on Sidekiq when due, exactly once. For long horizons (a job scheduled 30 days out) you may want to move them so that Redis can be retired sooner; do it with a move-and-delete that cannot duplicate.

# Move scheduled SendReceiptJob entries from Redis to Solid Queue, one at a time
Sidekiq::ScheduledSet.new.each do |entry|
  next unless entry.klass == "SendReceiptJob"
  ActiveRecord::Base.transaction do
    job = SendReceiptJob.new(*entry.args)
    job.scheduled_at = Time.at(entry.at)
    SolidQueue::Job.enqueue(job)                 # inserted, not yet committed
    raise ActiveRecord::Rollback unless entry.delete   # delete from Redis first-wins
  end
end

entry.delete returns false if Sidekiq already moved the job to a queue (it became due during the loop); rolling back the insert in that case leaves the job running on Sidekiq only. The reverse order — delete from Redis, then insert — risks losing the job if the process dies between the two.

Step 6 — Drain and Retire Redis

When every class is at 100% and the migration has run through at least one full cycle of your longest scheduled horizon, check that Redis holds nothing that still needs to run:

require "sidekiq/api"
puts "queued:    #{Sidekiq::Queue.all.sum(&:size)}"
puts "scheduled: #{Sidekiq::ScheduledSet.new.size}"
puts "retrying:  #{Sidekiq::RetrySet.new.size}"
puts "dead:      #{Sidekiq::DeadSet.new.size}"       # triage or export these first
puts "busy:      #{Sidekiq::Workers.new.size}"

All five must be zero (after exporting the dead set for triage). Scale the Sidekiq deployment to zero, keep Redis running read-only for a week in case a forgotten producer appears — its enqueue rate should stay at zero — then decommission it.

Migration timeline Week zero deploys the new workers idle. Weeks one to four ramp the four waves of job classes. Weeks five and six let scheduled and retrying jobs drain from Redis. Week seven scales Sidekiq to zero with Redis still available, and week eight decommissions Redis. Eight weeks, reversible until the last wk 0: idle wk 1-4: ramp waves 1 to 4 wk 5-6: drain wk 7 wk 8 Sidekiq to zero Redis gone Any class can be flipped back to Redis in minutes until week 8.

Verification

Reconcile counts end to end for each migrated class over a day:

# enqueued (from app metrics) == processed on Redis + processed on Postgres, per class
for cls in migrated_classes:
    enq = prom.query(f'sum(increase(jobs_enqueued_total{{class="{cls}"}}[1d]))')
    redis_done = prom.query(f'sum(increase(sidekiq_jobs_processed_total{{class="{cls}"}}[1d]))')
    pg_done = sql.scalar("SELECT count(*) FROM job_audit WHERE class=%s AND finished_at > now()-'1 day'::interval", cls)
    assert abs(enq - (redis_done + pg_done)) / max(enq, 1) < 0.001, cls

A mismatch above rounding error is a stop signal: pause the ramp for that class and find the missing or duplicated jobs before continuing.

Gotchas & Edge Cases

Serialization differences. Sidekiq stores JSON-native arguments; Active Job serializes GlobalID references. A class that passes a model instance behaves differently across adapters if the record is deleted before execution. Test deserialization failures on the target explicitly.

Queue priority semantics. Sidekiq's weighted queue polling and Solid Queue's ordered queue lists are not equivalent. Recreate the intended priority explicitly rather than copying queue names.

Hidden producers. Cron scripts, rake tasks, and other services may push directly to Redis with Sidekiq::Client.push, bypassing Active Job. The enqueue-rate check in Step 6 is what catches them.

Database load surprises. Wave 2 (high-volume classes) is where the database feels the change. Watch replication lag and autovacuum activity on queue tables as each class ramps, and be ready to pause.

Dashboards that follow the queue name, not the backend. During the migration a single logical queue is served by two systems. Alerts that look at only one backend's metrics will fire (or stay silent) for the wrong reasons as traffic shifts. Before the first flip, build a combined view per logical queue — enqueue rate from the application, processing rate and oldest-job age from both backends summed or maxed — and point alerts at that. Retire the per-backend alerts only after Redis is gone.

Retry counts reset on move. A job moved from the Sidekiq retry set in Step 5 starts on the new backend with a fresh attempt counter. For a job that has already failed eight times, that means another full retry budget. Carry the attempt count across in the job's arguments if your retry limits matter, or triage heavily retried jobs by hand instead of moving them.

FAQ

How long should each wave run before the next? At least one full business cycle — usually a week — so that daily and weekly peaks, scheduled batch jobs, and month-end spikes all hit the new backend at least once. Wave 4 (payment-critical classes) benefits from spanning a billing cycle.

Can I copy the whole Redis backlog to Postgres at once? You can, but it is the riskiest option: any bug in the copy duplicates or drops jobs at scale. Letting the backlog drain naturally on the old workers is slower and far safer.

How do I roll back a class? Flip its flag off. New jobs go to Redis again; jobs already in Postgres finish on the new workers. Because each job lives in one backend, rollback never duplicates work.

Does the same process work in reverse, from Postgres to Redis? Yes. Per-class enqueue routing, side-by-side fleets, and natural draining are backend-agnostic, which makes the same plan useful when a database queue outgrows the primary.

Related