Redis Sentinel for Queue High Availability

A single Redis node is a single point of failure for every queue that lives on it. Sentinel watches a primary and its replicas, promotes a replica when the primary dies, and tells clients where the new primary is. This guide sets it up for job queues and explains what a failover can and cannot protect, as part of In-Memory vs Persistent Queue Storage in Backend Frameworks & Worker Scaling.

Problem Statement

A self-hosted Redis node holds the queues for BullMQ and Celery. When the VM was live-migrated by the cloud provider, Redis was unreachable for four minutes; producers returned 500s, workers crashed in reconnect loops, and a batch of scheduled jobs fired late. A later kernel update required a planned restart that meant another outage window. The team wants automatic failover within about 30 seconds, clients that follow the primary without a redeploy, and a clear understanding of how many jobs can be lost when the primary fails.

Prerequisites

  • Three hosts (or availability zones) for Sentinel processes; two can also run Redis nodes.
  • Redis 6.2 or newer on a primary and at least one replica.
  • Client libraries that support Sentinel: ioredis (BullMQ), redis-py / kombu (Celery, RQ), redis-client (Sidekiq 7).
  • maxmemory-policy noeviction on every data node — see choosing a Redis maxmemory policy for queues.

Step 1 — Lay Out the Topology

Sentinel needs a majority to agree that the primary is down and to elect a leader to perform the failover. Three Sentinels tolerate one failure; place them in different failure domains so that a single host or zone outage cannot remove the majority.

Sentinel topology across three zones Zone A hosts the Redis primary and sentinel 1. Zone B hosts a replica, which receives asynchronous replication from the primary, and sentinel 2. Zone C hosts only sentinel 3, acting as a tie-breaker. Producers and workers first ask any sentinel for the current primary address, then connect to it directly. Primary, replica, and a Sentinel quorum of three zone A zone B zone C Redis primary Redis replica async sentinel 1 sentinel 2 sentinel 3 tie-breaker only clients: ask a sentinel, then connect to the primary

With two data nodes, losing the primary leaves one; add a second replica if you need to survive two node failures or want a replica to keep serving during maintenance.

Step 2 — Configure Sentinel

Each Sentinel gets the same monitor definition. quorum is how many Sentinels must agree the primary is unreachable before failover begins; a majority of all Sentinels must still authorise it.

# sentinel.conf
port 26379
sentinel monitor queues 10.0.1.10 6379 2
sentinel auth-pass queues ${REDIS_PASSWORD}
sentinel down-after-milliseconds queues 5000
sentinel failover-timeout queues 30000
sentinel parallel-syncs queues 1
sentinel resolve-hostnames yes

down-after-milliseconds trades detection speed against false failovers. Five seconds is reasonable on a stable network; below two seconds, a garbage-collection pause on the host or a brief network blip can trigger an unnecessary failover, which itself causes a short outage.

On the data nodes, limit how much data a failover can lose by refusing writes when replication is too far behind:

# redis.conf on primary and replicas
min-replicas-to-write 1
min-replicas-max-lag 10
appendonly yes
appendfsync everysec

With these settings, the primary rejects writes if no replica has acknowledged replication within 10 seconds, so a primary cut off from its replica cannot keep accepting jobs that will be discarded when the other side promotes a new primary.

Step 3 — Connect the Queue Libraries Through Sentinel

Clients connect to Sentinels, ask for the primary of the named master, and reconnect when Sentinel announces a switch. Never hard-code the primary's address.

BullMQ (ioredis):

const connection = new IORedis({
  sentinels: [
    { host: "sentinel-1", port: 26379 },
    { host: "sentinel-2", port: 26379 },
    { host: "sentinel-3", port: 26379 },
  ],
  name: "queues",
  password: process.env.REDIS_PASSWORD,
  maxRetriesPerRequest: null,   // required for BullMQ workers
  enableReadyCheck: true,
});
const worker = new Worker("invoices", processor, { connection });

Celery (kombu):

broker_url = "sentinel://:pw@sentinel-1:26379;sentinel://:pw@sentinel-2:26379;sentinel://:pw@sentinel-3:26379"
broker_transport_options = {"master_name": "queues", "visibility_timeout": 3600}
result_backend = broker_url
result_backend_transport_options = {"master_name": "queues"}

Sidekiq 7:

Sidekiq.configure_server do |config|
  config.redis = {
    sentinels: [{ host: "sentinel-1", port: 26379 }, { host: "sentinel-2", port: 26379 }],
    name: "queues", password: ENV["REDIS_PASSWORD"], role: :master
  }
end

Keep maxRetriesPerRequest: null for BullMQ workers so a failover produces a delay rather than an exception that stops the worker.

Producers need the opposite trade-off. A web request that enqueues a job should not hang for 15 seconds waiting for a new primary; give the producer connection a short command timeout and a bounded number of retries, and fall back to an outbox or an error response when they run out. Check the reconnect behaviour of every process type: Celery workers reconnect through kombu's retry policy, Sidekiq's redis-client re-resolves the primary on connection errors, and long-lived Node.js producers keep a single ioredis connection that follows +switch-master announcements automatically. Short-lived scripts and cron jobs that build their own connection are the ones most often left pointing at a fixed address.

Step 4 — Understand What a Failover Loses

Redis replication is asynchronous. When the primary crashes, any writes it acknowledged but had not yet sent to the replica are gone once the replica is promoted. For a queue, that means:

A failover timeline for a queue Before the crash, the replica trails the primary by a small replication lag, typically milliseconds. Writes inside that lag window at the moment of the crash are lost: enqueues that returned success, and state changes such as job completions. Then Sentinel waits down-after-milliseconds, elects a leader, promotes the replica, and clients reconnect, which takes roughly 5 to 15 seconds in total. From crash to new primary normal operation primary crashes lost: last lag window detect (5 s) promote reconnect writes fail or wait: roughly 5–15 s Lost writes can be enqueues that returned OK or completions a worker already reported.
  • Lost enqueues. A job acknowledged to the producer can disappear. If that is unacceptable, enqueue through a transactional outbox so the source of truth is your database and the relay can re-send after failover.
  • Lost completions. A job marked completed on the old primary can reappear as active or waiting on the new one and run again. Workers must be idempotent — which they must be anyway under at-least-once delivery.
  • Stalled jobs. Locks held by workers during the switch may expire; BullMQ's stalled-job checker moves those jobs back to waiting.

Redis offers WAIT numreplicas timeout to make a write wait until replicas have it. Most queue libraries do not use it, and it costs latency on every enqueue; the outbox approach is usually the better trade.

Step 5 — Handle Scheduled and Repeatable Jobs

Delayed and repeatable jobs live in sorted sets on the primary. A failover does not lose them unless they were written inside the lag window, but schedulers must reconnect promptly. Celery beat stores its schedule outside Redis by default and will simply keep publishing once the broker is back; BullMQ job schedulers are stored in Redis and are processed by whichever worker next checks the delayed set. After failover, check that the next scheduled run appears — see BullMQ job schedulers for repeatable jobs.

Step 6 — Test Failover Regularly

Practise in staging with production-like load before relying on it:

# trigger a planned failover
redis-cli -p 26379 SENTINEL FAILOVER queues

# or simulate a crash on the primary
redis-cli -h 10.0.1.10 DEBUG SLEEP 30

# watch the switch
redis-cli -p 26379 SENTINEL get-master-addr-by-name queues

While a load generator enqueues numbered jobs, record: time until enqueues succeed again, number of enqueue errors, number of jobs that ran twice, and number of acknowledged job IDs that never ran.

What to measure in a failover drill Four measurements from a failover drill, each with a target: recovery time under 30 seconds, enqueue errors retried by producers with none reaching users, duplicate runs harmless because handlers are idempotent, and missing jobs zero when enqueues go through an outbox. Failover drill scorecard recovery time < 30 s first successful enqueue enqueue errors retried none surface to users duplicate runs harmless idempotent handlers missing jobs 0 with an outbox Run the drill under load; an idle failover proves little.

Verification

  • SENTINEL masters on each Sentinel shows the same primary, num-other-sentinels 2, and quorum 2.
  • INFO replication on the primary shows the replica online with a lag of 0–1 seconds.
  • A forced failover completes in under 30 seconds and workers resume without a restart.
  • In the drill, no acknowledged job goes missing when producers use the outbox, and duplicate runs have no visible side effects.
  • Alerts fire on +switch-master events and on replica lag above a few seconds.

Gotchas & Edge Cases

Split brain after a network partition. If the old primary is isolated but still reachable by some clients, it can accept writes that are thrown away when it rejoins as a replica. min-replicas-to-write limits how long that can go on.

NAT and container networking. Sentinel announces the IP addresses it sees. Behind NAT or in Docker bridge networks, clients receive addresses they cannot reach. Use replica-announce-ip and sentinel announce-ip, or host networking.

Sentinel is not sharding. It gives high availability for one data set; it does not spread load. If a single primary cannot hold your queues, look at running BullMQ on Redis Cluster or split queues across instances.

Managed services. ElastiCache, Memorystore, and Azure Cache provide their own failover and a stable endpoint; you connect to that endpoint, not to Sentinel. The data-loss analysis in Step 4 still applies.

FAQ

How many Sentinels do I need? At least three, on independent hosts. Two Sentinels cannot form a majority if one fails, so failover never happens.

Does Sentinel make Redis as durable as a database? No. Asynchronous replication means a small window of acknowledged writes can be lost. Combine it with AOF persistence, idempotent handlers, and an outbox for jobs you cannot lose.

Can workers read from replicas to reduce load? Not for queue operations. Fetching and acknowledging jobs are writes and must go to the primary; replicas are useful for dashboards and read-only monitoring.

What happens to jobs that were running during the failover? The worker keeps executing the job in memory. When it tries to mark the job complete, the command waits or fails until the new primary is available, then succeeds on retry. If the lock expired meanwhile, another worker may have picked the job up as stalled, so the job can run twice.

Related