Configuring the BullMQ Rate Limiter
BullMQ ships a queue-wide rate limiter enforced in Redis, so the limit holds across every worker process no matter how many you run. This guide configures it for a rate-limited third-party API, including reacting to 429 responses at runtime, as part of Rate Limiting & Throttling Jobs in Queue Fundamentals & Architecture.
Problem Statement
A marketing service sends SMS through a provider that allows 50 requests per second per account and returns 429 with a Retry-After header when exceeded. The service runs 6 worker pods with concurrency 20 each. During campaigns, 120 concurrent jobs fire at once, the provider returns hundreds of 429s, BullMQ retries them with exponential backoff, and some messages arrive 40 minutes late while others exhaust their attempts. Setting concurrency to 50 total did not help, because concurrency limits jobs in flight, not requests per second. You want at most 50 SMS per second across the whole fleet, automatic slowdown when the provider signals a limit, and no failed attempts spent on throttling.
Prerequisites
- BullMQ 5.x (the
rateLimit/RateLimitErrorAPI described here is from BullMQ 3+). - All workers for the queue using the same queue name and Redis instance.
- The provider's documented limits and its throttling response format.
- Metrics for completed jobs per second and 429 responses.
Step 1 — Understand Concurrency vs Rate
Concurrency caps how many jobs a worker runs at once; the rate limiter caps how many jobs the queue starts per time window, across all workers. They solve different problems and are usually both needed.
concurrency = 20 per worker x 6 workers = 120 jobs in flight
job duration ≈ 200 ms -> up to 600 jobs/s could start
provider limit = 50 req/s
=> concurrency alone does not bound the rate; a limiter of 50 per 1000 ms does
A job that takes 200 ms with 120 slots can start 600 times a second. Only a limiter measured in jobs per duration matches a provider limit measured in requests per second.
Step 2 — Configure the Queue-Wide Limiter
The limiter is set on the worker options, but its state lives in Redis per queue, so every worker enforces the same shared budget. Give every worker for the queue the same limiter configuration.
import { Worker } from "bullmq";
const worker = new Worker("sms", sendSms, {
connection,
concurrency: 20, // memory/connection bound per process
limiter: {
max: 45, // stay a little under the provider's 50/s
duration: 1000, // per 1000 ms window
},
});
Leaving about 10% headroom absorbs clock differences between BullMQ's window and the provider's, plus any other clients sharing the same provider account. When the limit is reached, workers stop fetching until the window resets; waiting jobs stay in the queue and are not charged an attempt.
Step 3 — React to 429 Responses with Manual Rate Limiting
Even with a limiter, providers throttle for their own reasons (account-wide limits, burst rules, incidents). When the provider returns 429, tell BullMQ to pause the whole queue for the Retry-After period, and hand the job back without counting a failed attempt.
import { Worker, RateLimitError } from "bullmq";
async function sendSms(job: Job<SmsData>, token?: string) {
const res = await provider.send(job.data.to, job.data.body, { idempotencyKey: job.id });
if (res.status === 429) {
const retryAfterMs = Number(res.headers["retry-after"] ?? 1) * 1000;
await worker.rateLimit(retryAfterMs); // pause fetching queue-wide
throw new RateLimitError(); // job returns to waiting; no attempt used
}
if (res.status >= 500) throw new Error(`provider ${res.status}`); // normal retry path
if (res.status >= 400) throw new UnrecoverableError(`rejected: ${res.body.code}`);
return { messageId: res.body.id };
}
worker.rateLimit(ms) sets a queue-wide pause visible to every worker, so one 429 slows the whole fleet rather than each worker discovering the limit separately. Throwing RateLimitError moves the job back to waiting with its attempt count unchanged — the fix for jobs exhausting attempts on throttling in the problem statement.
Step 4 — Separate Limits per Account or Tenant
The limiter applies to the whole queue. When limits are per provider account (each customer has its own SMS sender with its own quota), a single queue would let one busy account consume everyone's budget. Two common approaches:
// A. One queue per account, each with its own limiter (good for tens of accounts)
function workerFor(accountId: string, max: number) {
return new Worker(`sms:${accountId}`, sendSms, { connection, concurrency: 5,
limiter: { max, duration: 1000 } });
}
// B. BullMQ Pro groups: one queue, per-group rate limits and fair rotation (many accounts)
await queue.add("send", data, { group: { id: accountId } });
// worker: new WorkerPro("sms", sendSms, { connection, group: { limit: { max: 10, duration: 1000 } } })
Per-queue limiters are available in open-source BullMQ and work well for a modest number of accounts. For thousands of tenants, BullMQ Pro's groups (or a custom token bucket per tenant, as in sliding window rate limiting with Redis Lua) avoid one queue per tenant. The fairness side of the problem is covered in preventing tenant starvation with weighted queues.
Step 5 — Monitor the Limiter Instead of Guessing
A rate-limited queue always has a backlog during bursts; the questions are whether it drains in acceptable time and whether 429s are rare.
const events = new QueueEvents("sms", { connection });
let completed = 0;
events.on("completed", () => completed++);
setInterval(async () => {
const counts = await queue.getJobCounts("waiting", "delayed", "active");
const ttl = await queue.getRateLimitTtl(); // ms until the limiter window/pause ends
smsCompletedPerSec.set(completed / 10); completed = 0;
smsWaiting.set(counts.waiting);
smsRateLimitTtl.set(ttl);
}, 10_000);
# Throughput should sit at the limit during campaigns, not below it
avg_over_time(sms_completed_per_sec[5m])
# 429s should be rare after the limiter is tuned
sum(rate(sms_provider_responses_total{status="429"}[5m])) / sum(rate(sms_provider_responses_total[5m])) > 0.01
# Backlog drain estimate: waiting / limit
sms_waiting / 45
If completions sit well below the limit while jobs wait, workers are the bottleneck (too little concurrency for the job duration); if 429s stay high, lower max or check for other clients sharing the account.
Verification
it("never exceeds the configured rate across two workers", async () => {
const stamps: number[] = [];
const handler = async () => { stamps.push(Date.now()); };
const w1 = new Worker(q.name, handler, { connection, concurrency: 50, limiter: { max: 10, duration: 1000 } });
const w2 = new Worker(q.name, handler, { connection, concurrency: 50, limiter: { max: 10, duration: 1000 } });
await q.addBulk(Array.from({ length: 50 }, () => ({ name: "x", data: {} })));
await waitUntil(() => stamps.length === 50, 10_000);
for (const t of stamps) {
const inWindow = stamps.filter((s) => s >= t && s < t + 1000).length;
expect(inWindow).toBeLessThanOrEqual(10);
}
});
Gotchas & Edge Cases
Different limiter settings per worker. The limiter state is shared, but each worker applies its own max. Deploy the same configuration everywhere, or the effective limit depends on which worker fetches.
Delayed jobs and the limiter. Jobs delayed by backoff re-enter waiting and compete for the same budget; a large retry wave can crowd out fresh jobs. Keep retry volume low by fixing the causes of failure.
Priority under rate limits. Priorities decide which job starts next within the limit; they do not raise the limit. Urgent messages still wait for their turn in the window.
Rate limiting and graceful shutdown. A worker closing while the queue is rate-limited waits for in-flight jobs only; jobs waiting for the window are untouched and will be picked up by the remaining workers. There is no need to drain the limiter before deploys.
Idempotency on 429 retries. Some providers return 429 after partially accepting a request. Always send an idempotency key (the job id works) so a retried send after a throttle cannot produce a duplicate SMS.
Provider windows differ. A provider enforcing a rolling window may still throttle a fixed-window limiter at window boundaries. Headroom and the 429 handler cover it.
FAQ
Does the limiter count failed jobs?
Every job start counts, including ones that fail. Jobs returned with RateLimitError count against the window they were started in.
How long will a campaign take to send under the limit? Divide the number of messages by the effective rate: 90,000 SMS at 45 per second is about 33 minutes. Put that number in front of whoever schedules campaigns, because it is a property of the provider contract, not of the worker fleet — adding workers will not shorten it. If the business needs faster delivery, the lever is a higher provider limit or splitting traffic across accounts that each have their own quota.
Should the limit live in code or configuration? Configuration, read at worker start, so it can be adjusted when the provider changes a quota without a code deploy. Log the active limit on startup and export it as a metric, so dashboards show throughput against the limit actually in force.
Can I limit by job name within one queue? Not with the built-in limiter — it is per queue. Use separate queues per limit, or implement a token bucket keyed by name.
What's the difference between worker.rateLimit and pausing the queue?
rateLimit(ms) is a timed pause that expires on its own and is meant for throttling; queue.pause() stops processing until someone resumes it and is meant for incidents or maintenance.
Related
- Rate Limiting & Throttling Jobs — rate limiting patterns across frameworks.
- Rate Limiting Third-Party API Calls from Workers — framework-neutral approaches.
- Configuring BullMQ Concurrency Limits for High Throughput — the other half of the configuration.
- BullMQ Retries and Backoff Strategies — what happens on real failures.