BullMQ Failed Jobs as a Dead-Letter Queue
BullMQ does not have a component called a dead-letter queue, but every queue already has one: the failed set, where jobs land once they exhaust their attempts. This guide shows how to run that set deliberately — retention, inspection, alerting, and replay — as part of Dead-Letter Queues & Poison Messages in Queue Fundamentals & Architecture.
Problem Statement
An order-sync service processes BullMQ jobs that push orders to an ERP system. Workers are configured with removeOnFail: true to "keep Redis clean", so when an ERP outage caused 8,000 jobs to exhaust their retries, they vanished without a trace — the orders never reached the ERP, and finance discovered the gap during month-end reconciliation. On another queue the opposite happened: removeOnFail was unset, and 2 million failed jobs accumulated over a year, consuming 6 GB of Redis memory. You want failed jobs kept long enough to triage and replay, bounded so they cannot fill Redis, alerting when they appear, and a safe bulk-retry procedure after a fix.
Prerequisites
- BullMQ 5.x workers and producers, a Redis instance with
noeviction. - The ability to change job options (
attempts,backoff,removeOnFail) at enqueue or as queue defaults. - A place to run operational scripts (a small admin CLI or a one-off job).
- Metrics collection for queue counts (BullMQ's
getJobCounts, a Prometheus exporter, or Bull Board).
Step 1 — Make Jobs Reach the Failed Set for the Right Reasons
A job enters the failed set when it throws on its final attempt, or immediately if it throws UnrecoverableError. Configure attempts and backoff for transient failures, and fail fast on permanent ones so the failed set contains only what needs attention.
import { Queue, Worker, UnrecoverableError } from "bullmq";
const queue = new Queue("erp-sync", {
connection,
defaultJobOptions: {
attempts: 8,
backoff: { type: "exponential", delay: 30_000 }, // 30 s, 60 s, 2 min ... ~1 h total
removeOnComplete: { age: 3600, count: 1000 },
removeOnFail: { age: 14 * 24 * 3600, count: 50_000 }, // Step 2
},
});
const worker = new Worker("erp-sync", async (job) => {
const order = await loadOrder(job.data.orderId);
if (!order) throw new UnrecoverableError(`order ${job.data.orderId} deleted`); // straight to failed
const res = await erp.upsertOrder(order, { idempotencyKey: `order-${order.id}` });
if (res.status === 422) throw new UnrecoverableError(`ERP rejected: ${res.body.error}`);
if (res.status >= 500 || res.status === 429) throw new Error(`ERP ${res.status}`); // retry
}, { connection, concurrency: 20 });
Eight attempts with exponential backoff span about an hour; an ERP outage shorter than that recovers without anyone touching the failed set. A longer outage fills it — which is exactly when you want the jobs preserved rather than deleted.
Step 2 — Bound Retention by Age and Count
removeOnFail accepts true (delete immediately — no dead letters at all), false (keep forever — unbounded memory), a number (keep the last N), or { age, count } (keep up to N jobs no older than age seconds). The last form is the one to use.
removeOnFail: {
age: 14 * 24 * 3600, // 14 days: covers triage, weekends, and a fix-and-deploy cycle
count: 50_000, // hard cap: bounds memory even during a mass-failure event
}
Size the count from memory: a failed job stores its data, options, stack trace, and attempt history. Measure with MEMORY USAGE bull:erp-sync:<jobId> on a few samples — typically 2–10 KB each — so 50,000 failed jobs cost 100–500 MB. If a mass failure could exceed the cap, the oldest failures are removed first; Step 4 moves them somewhere durable before that happens. Memory planning for queues is covered in sizing Redis memory for queue backlogs.
Step 3 — Inspect Failures by Reason
Before retrying anything, group failures by reason to see whether you are looking at one cause or many.
// scripts/failed-report.ts
const failed = await queue.getFailed(0, 4999); // newest first
const byReason = new Map<string, { count: number; sample: string }>();
for (const job of failed) {
const key = (job.failedReason ?? "unknown").replace(/\d+/g, "N").slice(0, 120);
const e = byReason.get(key) ?? { count: 0, sample: job.id! };
e.count++; byReason.set(key, e);
}
console.table([...byReason].sort((a, b) => b[1].count - a[1].count)
.map(([reason, v]) => ({ reason, count: v.count, sampleJob: v.sample })));
// ERP 503 7_912 job 88121
// ERP rejected: invalid tax code 61 job 90112
// order N deleted 27 job 90555
Normalising digits collapses messages that differ only in ids. The report above says: retry the 503s once the ERP is healthy; fix the tax-code mapping before retrying those 61; discard the deleted-order jobs. Bull Board or Taskforce provide the same view interactively.
Step 4 — Retry in Bulk, Safely
Retry by reason, in batches, with a rate limit, and only after confirming the cause is fixed. job.retry() moves a failed job back to waiting with its attempts reset.
// scripts/retry-failed.ts --reason "ERP 503" --batch 500 --pause-ms 2000
async function retryByReason(match: RegExp, batch = 500, pauseMs = 2000, dryRun = true) {
let retried = 0;
for (let start = 0; ; start += batch) {
const jobs = await queue.getFailed(start, start + batch - 1);
if (jobs.length === 0) break;
const selected = jobs.filter((j) => match.test(j.failedReason ?? ""));
if (!dryRun) await Promise.all(selected.map((j) => j.retry("failed")));
retried += selected.length;
await new Promise((r) => setTimeout(r, pauseMs)); // don't flood a recovering ERP
}
console.log(`${dryRun ? "would retry" : "retried"} ${retried} jobs`);
}
Paging with getFailed while retrying shifts indices (retried jobs leave the set), so production scripts should either collect ids first and then retry them, or loop until no matching jobs remain. For "retry everything", BullMQ also offers queue.retryJobs({ state: "failed", count: 1000 }), which does the move server-side in batches. Replay discipline — fix first, validate on a few, then rate-limited bulk — is the same as in replaying dead-letter messages in RabbitMQ.
Step 5 — Optionally Forward Failures to a Durable Store
The failed set lives in Redis with bounded retention. For jobs that must never be lost (financial syncs), forward each final failure to a durable store as it happens, so retention limits and Redis incidents cannot erase evidence.
const events = new QueueEvents("erp-sync", { connection });
events.on("failed", async ({ jobId, failedReason, prev }) => {
const job = await Job.fromId(queue, jobId);
if (!job || job.attemptsMade < (job.opts.attempts ?? 1)) return; // not final yet
await db.query(
`INSERT INTO dead_letters (queue, job_id, name, data, reason, attempts, failed_at)
VALUES ($1,$2,$3,$4,$5,$6, now()) ON CONFLICT (queue, job_id) DO NOTHING`,
["erp-sync", jobId, job.name, job.data, failedReason, job.attemptsMade]);
});
The database record survives Redis eviction, flushes, and the removeOnFail limits, and it is easy to query in reconciliation. Replaying from it means enqueueing a new job with the stored data.
Step 6 — Alert on Failed-Set Growth
A failed set that nobody watches is how 8,000 orders disappeared. Alert on growth rate and on absolute size, per queue.
# New failures in the last 10 minutes (from a counter incremented on final failure)
sum by (queue) (increase(bullmq_jobs_failed_final_total[10m])) > 20
# Failed-set size approaching its retention cap (from getJobCounts exported as a gauge)
bullmq_queue_jobs{state="failed"} / 50000 > 0.5
The first alert catches incidents in progress; the second catches slow accumulation and warns before retention starts deleting evidence. Routing and thresholds follow alerting on dead-letter queue growth.
Verification
it("keeps failed jobs with reasons and retries them after a fix", async () => {
erpFake.failWith(503);
const job = await queue.add("sync", { orderId: "o-1" }, { attempts: 2, backoff: { type: "fixed", delay: 50 } });
await expect(job.waitUntilFinished(events, 5000)).rejects.toThrow(/ERP 503/);
expect((await queue.getJobCounts("failed")).failed).toBe(1);
erpFake.recover();
await (await queue.getJob(job.id!))!.retry("failed");
await expect((await queue.getJob(job.id!))!.waitUntilFinished(events, 5000)).resolves.toBeUndefined();
});
Gotchas & Edge Cases
removeOnFail: true in examples. Many tutorials set it to keep Redis tidy. In production it deletes your dead-letter queue.
Retrying stale jobs. A job failed two weeks ago may no longer be valid (the order was edited or cancelled). Handlers should re-read current state, not trust job data blindly.
Flows and parent jobs. A failed child can leave its parent waiting; decide per flow whether children use failParentOnFailure or ignoreDependencyOnFailure.
Retention cap during mass failure. If failures exceed count, the oldest are removed. Forward to durable storage (Step 5) for anything critical.
FAQ
Should I use a separate "dlq" queue instead of the failed set?
Only if you need different tooling or retention for dead letters. The failed set already stores reasons, stack traces, and attempts, and retry() works in place; a separate queue adds a move step without much benefit.
How do I stop an outage from filling the failed set in the first place?
Pause the queue when a dependency is known to be down (queue.pause(), or a circuit breaker in the processor that throws a retryable error with a long delay), so jobs wait instead of burning through their attempts. A failed set full of identical 503s is usually a sign that retries ran out before the outage did — lengthen the backoff horizon for dependencies with a history of long outages.
Do retried jobs keep their history?
retry() resets attemptsMade and moves the job back to waiting; the failed reason and stack trace from previous attempts remain in job.stacktrace.
How do I discard jobs that should not be retried?
job.remove() for individual jobs, or queue.clean(grace, limit, "failed") for bulk removal by age. Record what you discard and why.
Related
- Dead-Letter Queues & Poison Messages — DLQ principles across brokers.
- BullMQ Retries and Backoff Strategies — the attempts before a job fails.
- BullMQ for Node.js Ecosystems — production BullMQ setup.
- Handling Poison Messages in Kafka — the same problem on a log.