Redis Sentinel for Queue High Availability
A single Redis node is a single point of failure for every queue that lives on it. Sentinel watches a primary and its replicas, promotes a replica when the primary dies, and tells clients where the new primary is. This guide sets it up for job queues and explains what a failover can and cannot protect, as part of In-Memory vs Persistent Queue Storage in Backend Frameworks & Worker Scaling.
Problem Statement
A self-hosted Redis node holds the queues for BullMQ and Celery. When the VM was live-migrated by the cloud provider, Redis was unreachable for four minutes; producers returned 500s, workers crashed in reconnect loops, and a batch of scheduled jobs fired late. A later kernel update required a planned restart that meant another outage window. The team wants automatic failover within about 30 seconds, clients that follow the primary without a redeploy, and a clear understanding of how many jobs can be lost when the primary fails.
Prerequisites
- Three hosts (or availability zones) for Sentinel processes; two can also run Redis nodes.
- Redis 6.2 or newer on a primary and at least one replica.
- Client libraries that support Sentinel: ioredis (BullMQ), redis-py / kombu (Celery, RQ), redis-client (Sidekiq 7).
maxmemory-policy noevictionon every data node — see choosing a Redis maxmemory policy for queues.
Step 1 — Lay Out the Topology
Sentinel needs a majority to agree that the primary is down and to elect a leader to perform the failover. Three Sentinels tolerate one failure; place them in different failure domains so that a single host or zone outage cannot remove the majority.
With two data nodes, losing the primary leaves one; add a second replica if you need to survive two node failures or want a replica to keep serving during maintenance.
Step 2 — Configure Sentinel
Each Sentinel gets the same monitor definition. quorum is how many Sentinels must agree the primary is unreachable before failover begins; a majority of all Sentinels must still authorise it.
# sentinel.conf
port 26379
sentinel monitor queues 10.0.1.10 6379 2
sentinel auth-pass queues ${REDIS_PASSWORD}
sentinel down-after-milliseconds queues 5000
sentinel failover-timeout queues 30000
sentinel parallel-syncs queues 1
sentinel resolve-hostnames yes
down-after-milliseconds trades detection speed against false failovers. Five seconds is reasonable on a stable network; below two seconds, a garbage-collection pause on the host or a brief network blip can trigger an unnecessary failover, which itself causes a short outage.
On the data nodes, limit how much data a failover can lose by refusing writes when replication is too far behind:
# redis.conf on primary and replicas
min-replicas-to-write 1
min-replicas-max-lag 10
appendonly yes
appendfsync everysec
With these settings, the primary rejects writes if no replica has acknowledged replication within 10 seconds, so a primary cut off from its replica cannot keep accepting jobs that will be discarded when the other side promotes a new primary.
Step 3 — Connect the Queue Libraries Through Sentinel
Clients connect to Sentinels, ask for the primary of the named master, and reconnect when Sentinel announces a switch. Never hard-code the primary's address.
BullMQ (ioredis):
const connection = new IORedis({
sentinels: [
{ host: "sentinel-1", port: 26379 },
{ host: "sentinel-2", port: 26379 },
{ host: "sentinel-3", port: 26379 },
],
name: "queues",
password: process.env.REDIS_PASSWORD,
maxRetriesPerRequest: null, // required for BullMQ workers
enableReadyCheck: true,
});
const worker = new Worker("invoices", processor, { connection });
Celery (kombu):
broker_url = "sentinel://:pw@sentinel-1:26379;sentinel://:pw@sentinel-2:26379;sentinel://:pw@sentinel-3:26379"
broker_transport_options = {"master_name": "queues", "visibility_timeout": 3600}
result_backend = broker_url
result_backend_transport_options = {"master_name": "queues"}
Sidekiq 7:
Sidekiq.configure_server do |config|
config.redis = {
sentinels: [{ host: "sentinel-1", port: 26379 }, { host: "sentinel-2", port: 26379 }],
name: "queues", password: ENV["REDIS_PASSWORD"], role: :master
}
end
Keep maxRetriesPerRequest: null for BullMQ workers so a failover produces a delay rather than an exception that stops the worker.
Producers need the opposite trade-off. A web request that enqueues a job should not hang for 15 seconds waiting for a new primary; give the producer connection a short command timeout and a bounded number of retries, and fall back to an outbox or an error response when they run out. Check the reconnect behaviour of every process type: Celery workers reconnect through kombu's retry policy, Sidekiq's redis-client re-resolves the primary on connection errors, and long-lived Node.js producers keep a single ioredis connection that follows +switch-master announcements automatically. Short-lived scripts and cron jobs that build their own connection are the ones most often left pointing at a fixed address.
Step 4 — Understand What a Failover Loses
Redis replication is asynchronous. When the primary crashes, any writes it acknowledged but had not yet sent to the replica are gone once the replica is promoted. For a queue, that means:
- Lost enqueues. A job acknowledged to the producer can disappear. If that is unacceptable, enqueue through a transactional outbox so the source of truth is your database and the relay can re-send after failover.
- Lost completions. A job marked completed on the old primary can reappear as active or waiting on the new one and run again. Workers must be idempotent — which they must be anyway under at-least-once delivery.
- Stalled jobs. Locks held by workers during the switch may expire; BullMQ's stalled-job checker moves those jobs back to waiting.
Redis offers WAIT numreplicas timeout to make a write wait until replicas have it. Most queue libraries do not use it, and it costs latency on every enqueue; the outbox approach is usually the better trade.
Step 5 — Handle Scheduled and Repeatable Jobs
Delayed and repeatable jobs live in sorted sets on the primary. A failover does not lose them unless they were written inside the lag window, but schedulers must reconnect promptly. Celery beat stores its schedule outside Redis by default and will simply keep publishing once the broker is back; BullMQ job schedulers are stored in Redis and are processed by whichever worker next checks the delayed set. After failover, check that the next scheduled run appears — see BullMQ job schedulers for repeatable jobs.
Step 6 — Test Failover Regularly
Practise in staging with production-like load before relying on it:
# trigger a planned failover
redis-cli -p 26379 SENTINEL FAILOVER queues
# or simulate a crash on the primary
redis-cli -h 10.0.1.10 DEBUG SLEEP 30
# watch the switch
redis-cli -p 26379 SENTINEL get-master-addr-by-name queues
While a load generator enqueues numbered jobs, record: time until enqueues succeed again, number of enqueue errors, number of jobs that ran twice, and number of acknowledged job IDs that never ran.
Verification
SENTINEL masterson each Sentinel shows the same primary,num-other-sentinels 2, andquorum 2.INFO replicationon the primary shows the replicaonlinewith a lag of 0–1 seconds.- A forced failover completes in under 30 seconds and workers resume without a restart.
- In the drill, no acknowledged job goes missing when producers use the outbox, and duplicate runs have no visible side effects.
- Alerts fire on
+switch-masterevents and on replica lag above a few seconds.
Gotchas & Edge Cases
Split brain after a network partition. If the old primary is isolated but still reachable by some clients, it can accept writes that are thrown away when it rejoins as a replica. min-replicas-to-write limits how long that can go on.
NAT and container networking. Sentinel announces the IP addresses it sees. Behind NAT or in Docker bridge networks, clients receive addresses they cannot reach. Use replica-announce-ip and sentinel announce-ip, or host networking.
Sentinel is not sharding. It gives high availability for one data set; it does not spread load. If a single primary cannot hold your queues, look at running BullMQ on Redis Cluster or split queues across instances.
Managed services. ElastiCache, Memorystore, and Azure Cache provide their own failover and a stable endpoint; you connect to that endpoint, not to Sentinel. The data-loss analysis in Step 4 still applies.
FAQ
How many Sentinels do I need? At least three, on independent hosts. Two Sentinels cannot form a majority if one fails, so failover never happens.
Does Sentinel make Redis as durable as a database? No. Asynchronous replication means a small window of acknowledged writes can be lost. Combine it with AOF persistence, idempotent handlers, and an outbox for jobs you cannot lose.
Can workers read from replicas to reduce load? Not for queue operations. Fetching and acknowledging jobs are writes and must go to the primary; replicas are useful for dashboards and read-only monitoring.
What happens to jobs that were running during the failover? The worker keeps executing the job in memory. When it tries to mark the job complete, the command waits or fails until the new primary is available, then succeeds on retry. If the lock expired meanwhile, another worker may have picked the job up as stalled, so the job can run twice.
Related
- In-Memory vs Persistent Queue Storage — durability trade-offs.
- Redis Persistence: AOF vs RDB for Queues — surviving restarts.
- Choosing a Redis maxmemory Policy for Queues — never evicting jobs.
- Running BullMQ on Redis Cluster — when one primary is not enough.