BullMQ Retries and Backoff Strategies
BullMQ retries a job when its processor throws, according to the job's attempts and backoff options. The defaults are easy to get almost right, and "almost" is where retry storms and silent data loss come from. This guide configures retries for real workloads as part of BullMQ for Node.js Ecosystems in Backend Frameworks & Worker Scaling.
Problem Statement
A Node.js integration service syncs CRM records through a partner API. Jobs are added with attempts: 3 and backoff: { type: "exponential", delay: 1000 }. During a 10-minute partner outage, every job exhausted its three attempts in about 7 seconds and landed in the failed set; 12,000 syncs had to be retried by hand. Validation errors from bad CRM data were retried three times each for no reason, and 429 responses with a Retry-After of 60 seconds were retried after 1, 2, and 4 seconds, earning more 429s. You want retries that outlast realistic outages, different behaviour per error class, respect for Retry-After, and visibility into retry volume.
Prerequisites
- BullMQ 5.x workers and producers.
- A classification of the errors your processor can throw (transient, rate-limited, permanent).
- Default job options set in one place (queue
defaultJobOptionsor a shared enqueue helper). - Metrics on job outcomes, or
QueueEventsaccess to build them.
Step 1 ā Size Attempts and Delay to Cover Real Outages
Total retry time is what matters, not the attempt count. With exponential backoff, BullMQ waits delay Ć 2^(attempt ā 1) before each retry, so the total time across n attempts is roughly delay Ć (2^(nā1) ā 1).
function totalRetrySpan(delayMs: number, attempts: number) {
let t = 0;
for (let a = 1; a < attempts; a++) t += delayMs * 2 ** (a - 1);
return t / 1000;
}
console.log(totalRetrySpan(1000, 3)); // 3 s -> the scenario: gone before any outage ends
console.log(totalRetrySpan(5000, 10)); // 2,555 s -> ~43 minutes of retrying
console.log(totalRetrySpan(10000, 12)); // 20,470 s -> ~5.7 hours
Choose the span from the longest outage you want to ride out without human action ā typically tens of minutes to a few hours for third-party APIs ā then set attempts and delay to reach it.
const queue = new Queue("crm-sync", {
connection,
defaultJobOptions: {
attempts: 10,
backoff: { type: "exponential", delay: 5_000, jitter: 0.5 }, // ~43 min span, jittered
removeOnFail: { age: 14 * 24 * 3600, count: 50_000 },
},
});
jitter (BullMQ 5.x) randomises each delay by up to the given fraction, spreading retries of jobs that failed at the same moment ā the defence against synchronised retry waves described in preventing retry storms after an outage.
Step 2 ā Fail Fast on Permanent Errors
Retrying a job that can never succeed wastes attempts and delays the alert. Throw UnrecoverableError for permanent failures; BullMQ moves the job straight to failed regardless of remaining attempts.
import { Worker, UnrecoverableError } from "bullmq";
const worker = new Worker("crm-sync", async (job) => {
const record = await crm.get(job.data.recordId);
if (!record) throw new UnrecoverableError(`record ${job.data.recordId} no longer exists`);
const res = await partner.upsert(mapRecord(record));
if (res.status === 400 || res.status === 422) {
throw new UnrecoverableError(`partner rejected: ${res.body?.error ?? res.status}`);
}
if (res.status >= 500) throw new TransientError(`partner ${res.status}`);
return { partnerId: res.body.id };
}, { connection, concurrency: 25 });
The classification mirrors retrying only transient errors by exception type: data problems fail immediately with a clear reason; infrastructure problems retry.
Step 3 ā Use a Custom Backoff Strategy per Error
A single backoff curve cannot fit both "server error, try again soon" and "rate limited, the server told you exactly when". Custom backoff strategies receive the attempt number and the error, and return a delay (or -1 to stop retrying).
class RateLimitedError extends Error {
constructor(public retryAfterMs: number) { super(`rate limited for ${retryAfterMs} ms`); }
}
const worker = new Worker("crm-sync", processor, {
connection,
settings: {
backoffStrategy: (attemptsMade: number, type: string, err: Error) => {
if (err instanceof RateLimitedError) return err.retryAfterMs + Math.random() * 2000;
if (type === "crm") {
const base = Math.min(5_000 * 2 ** (attemptsMade - 1), 30 * 60_000); // cap at 30 min
return Math.random() * base; // full jitter
}
return -1; // unknown type: no retry
},
},
});
// jobs opt into the custom strategy by type
await queue.add("sync", { recordId }, { attempts: 12, backoff: { type: "crm" } });
Honouring Retry-After exactly fixes the 429 loop in the problem statement. When the whole queue should slow down on a 429 ā not just the one job ā use the worker-level rate limit instead, as in configuring the BullMQ rate limiter.
Step 4 ā Delay a Job Without Counting an Attempt
Some situations are not failures: a record is locked by another sync, a dependency's circuit breaker is open, or data is not ready yet. Throwing would consume an attempt. Instead, move the job back to delayed and signal BullMQ not to treat it as a failure.
import { DelayedError } from "bullmq";
const worker = new Worker("crm-sync", async (job, token) => {
if (await breaker.isOpen("partner")) {
await job.moveToDelayed(Date.now() + 30_000 + Math.random() * 15_000, token);
throw new DelayedError(); // tells BullMQ: not a failure, no attempt used
}
// ... normal processing ...
}, { connection });
This is the building block for dependency-aware pausing; the breaker itself is described in circuit breakers for worker dependencies.
Step 5 ā Observe Retries, Not Just Failures
A queue that eventually succeeds can still be unhealthy if most jobs need several attempts. Track attempts at completion and retry events.
const events = new QueueEvents("crm-sync", { connection });
worker.on("completed", (job) =>
attemptsAtSuccess.observe({ queue: job.queueName }, job.attemptsMade + 1));
worker.on("failed", (job, err) => {
if (!job) return;
const final = job.attemptsMade >= (job.opts.attempts ?? 1) || err instanceof UnrecoverableError;
jobFailures.inc({ queue: job.queueName, final: String(final), error: err.name });
});
# Share of successful jobs that needed retries: rising values mean a degrading dependency
sum(rate(job_attempts_at_success_bucket{le="1"}[15m])) / sum(rate(job_attempts_at_success_count[15m]))
# Final failures by error type
sum by (error) (rate(job_failures_total{final="true"}[15m]))
A drop in first-attempt success is often the earliest sign of a dependency problem, well before anything reaches the failed set.
Step 6 ā Write Retry Settings Down per Job Type
Retry settings encode business decisions ā how long a sync may lag, how much a stale email is worth ā so they deserve a table that product and engineering agree on, rather than values scattered across enqueue calls.
| Job type | Attempts | Backoff | Span | Why |
|---|---|---|---|---|
| CRM sync | 12 | custom crm, 5 s base, full jitter, 30 min cap |
~5 h | partner outages last hours; data must arrive |
| Password reset email | 4 | exponential, 2 s | ~15 s | useless after a few minutes; user can request again |
| Invoice PDF | 8 | exponential, 10 s | ~20 min | user waits in-app; after that, alert and fix |
| Nightly export | 3 | fixed, 10 min | ~20 min | next night's run supersedes it |
Encode the table once, in the enqueue helper, keyed by job name, and have code review treat changes to it like any other contract change. The broader reasoning about budgets is in setting retry budgets and max attempts.
const RETRY: Record<string, Partial<JobsOptions>> = {
"crm-sync": { attempts: 12, backoff: { type: "crm" } },
"password-reset": { attempts: 4, backoff: { type: "exponential", delay: 2_000 } },
"invoice-pdf": { attempts: 8, backoff: { type: "exponential", delay: 10_000, jitter: 0.5 } },
"nightly-export": { attempts: 3, backoff: { type: "fixed", delay: 600_000 } },
};
export const enqueue = (q: Queue, name: string, data: unknown, opts: JobsOptions = {}) =>
q.add(name, data, { ...RETRY[name], ...opts });
Verification
it("retries transient errors and fails permanent ones immediately", async () => {
partnerFake.failTimes(2, 503);
const ok = await queue.add("sync", { recordId: "r1" }, { attempts: 5, backoff: { type: "fixed", delay: 10 } });
await ok.waitUntilFinished(events, 5000);
expect((await queue.getJob(ok.id!))!.attemptsMade).toBe(3);
partnerFake.respond(422);
const bad = await queue.add("sync", { recordId: "r2" }, { attempts: 5 });
await expect(bad.waitUntilFinished(events, 5000)).rejects.toThrow(/partner rejected/);
expect((await queue.getJob(bad.id!))!.attemptsMade).toBe(1);
});
The integration harness is described in integration testing BullMQ workers with Testcontainers.
Gotchas & Edge Cases
Default attempts is 1. Without attempts, a job never retries. Set it in defaultJobOptions, not per call site.
Backoff in seconds by mistake. delay is milliseconds. delay: 5 retries after 5 ms.
Retries and ordering. A retried job goes back behind newer jobs; code that assumes jobs for one entity run in order breaks under retries. See Message Ordering Guarantees.
Stalls are separate. Jobs that stall (lost lock) follow maxStalledCount, not attempts.
FAQ
Should every queue retry for hours? No ā match the span to the job's value over time. A login code is useless after five minutes; an invoice sync matters for days.
Can I change retry options for jobs already queued? Options are stored per job at enqueue time. New defaults apply to new jobs; queued jobs keep their options.
How do I retry a job manually after it failed?
Call job.retry() on the failed job, or queue.retryJobs({ state: "failed" }) in bulk once the cause is fixed. The attempt count resets, so the job gets its full retry budget again. Retry by failure reason rather than blindly, so jobs that failed for a still-unfixed reason do not go straight back to the failed set.
Should the processor catch errors and retry internally? Only for very short, local retries (a single immediate retry on a connection reset). Anything longer belongs to BullMQ's backoff, so the job releases its worker slot while it waits and the retry is visible in metrics and the UI.
What's the difference between exponential and fixed backoff? Fixed waits the same delay each time ā fine for short, predictable hiccups. Exponential grows the delay, which suits outages of unknown length.
Related
- BullMQ for Node.js Ecosystems ā the overall BullMQ setup.
- BullMQ Failed Jobs as a Dead-Letter Queue ā where exhausted jobs go.
- Retry Strategies & Backoff ā the principles behind these settings.
- Deduplicating Jobs in BullMQ ā keeping retries from creating duplicates.