Preventing Retry Storms After an Outage
The most dangerous moment for a dependency is not the outage but the minute after it recovers, when every job that failed during the outage retries at once. This guide shapes that recovery so the backlog drains instead of causing a second outage, as part of Retry Strategies & Backoff in Queue Fundamentals & Architecture.
Problem Statement
An email-rendering service depends on a template API. The template API was down for 20 minutes; during that time 60,000 jobs failed and were scheduled for retry with exponential backoff. When the API came back, the retries — plus the backlog of new jobs, plus autoscaled workers that had grown from 10 to 80 pods on queue depth — hit it at roughly 30 times its normal request rate. It fell over again within 90 seconds. The cycle repeated three times before an engineer manually scaled workers down. You want recovery that ramps load up gradually, retry timing spread out rather than synchronised, a cap on how much load the queue can put on the dependency regardless of backlog, and a way to drop work that no longer matters.
Prerequisites
- Retry configuration you control (framework retry options or your own policy function).
- Worker autoscaling you can bound (KEDA, HPA, or manual scaling).
- A way to cap concurrency or rate for calls to the dependency (queue-level limits or a shared limiter).
- Knowledge of the dependency's safe throughput, or a way to measure it.
Step 1 — See How Retries Synchronise
Plain exponential backoff, applied to jobs that all failed within the same few minutes, produces waves: every job that failed at attempt 3 retries after the same delay. Adding a small random amount ("equal jitter" of ±10%) barely helps when thousands of jobs share an attempt number.
def plain(attempt): return min(3600, 5 * 2 ** attempt) # synchronised waves
def small_jitter(attempt): return plain(attempt) * random.uniform(0.9, 1.1) # waves, slightly blurred
def full_jitter(attempt): return random.uniform(0, min(3600, 5 * 2 ** attempt)) # spread across the range
Full jitter draws each delay uniformly from zero to the exponential ceiling, which spreads a cohort of retries evenly over the whole interval instead of stacking them at its end. It is the single most effective change against retry storms and costs one line.
Step 2 — Cap Load on the Dependency, Not Just Worker Count
Autoscaling on queue depth adds workers exactly when the dependency can least afford more callers. Bound the total concurrency (or rate) of calls to the dependency independently of worker count, so scaling workers adds capacity for other work but not extra pressure here.
# A fleet-wide semaphore for calls to the template API (Redis-backed)
SEM_KEY, SEM_LIMIT = "sem:template-api", 40 # dependency handles ~40 concurrent renders
def with_template_slot(fn):
token = acquire_semaphore(SEM_KEY, SEM_LIMIT, ttl=30) # returns None if all slots busy
if token is None:
raise ThrottledByUs() # defer the job; no attempt consumed
try:
return fn()
finally:
release_semaphore(SEM_KEY, token)
# And bound the autoscaler: more workers than the dependency can serve only add contention
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
spec:
minReplicaCount: 10
maxReplicaCount: 20 # was 80: 20 pods x 2 slots = 40 = dependency concurrency
Matching maximum concurrency to the dependency's capacity is the same principle as the Lambda concurrency cap in processing SQS with AWS Lambda. For rate-based limits use a limiter instead of a semaphore, as in sliding window rate limiting with Redis Lua.
Step 3 — Ramp Up Gradually After Recovery
Even a correct cap is too much at the instant of recovery, when the dependency's caches are cold and its own autoscaling has not caught up. Raise the cap in steps after the circuit breaker closes (or after errors stop).
RAMP = [5, 10, 20, 40] # concurrent slots, one step per minute of clean operation
def current_limit() -> int:
since = time.time() - float(r.get("template-api:recovered_at") or 0)
step = int(since // 60)
return RAMP[min(step, len(RAMP) - 1)]
def on_breaker_closed():
r.set("template-api:recovered_at", time.time())
# acquire_semaphore(SEM_KEY, current_limit(), ttl=30)
If errors reappear during the ramp, the breaker reopens and the ramp restarts from the bottom. This "slow start" is standard in load balancers for exactly this reason; it belongs in job systems too. Breakers are covered in circuit breakers for worker dependencies.
Step 4 — Enforce a Retry Budget
Retries should be a bounded share of traffic. A retry budget caps retries at, say, 20% of recent calls to a dependency; beyond that, failing jobs are deferred longer or parked instead of retried immediately.
def may_retry_now(dep: str, budget_ratio: float = 0.2) -> bool:
calls = int(r.get(f"calls:{dep}:1m") or 0)
retries = int(r.get(f"retries:{dep}:1m") or 0)
return retries < max(10, calls * budget_ratio)
def on_transient_failure(task, dep):
if may_retry_now(dep):
r.incr(f"retries:{dep}:1m"); r.expire(f"retries:{dep}:1m", 60)
raise task.retry(countdown=full_jitter(task.request.retries))
raise task.retry(countdown=300 + random.uniform(0, 300), max_retries=None) # park, don't pile on
The budget keeps an outage from turning the queue's traffic into mostly retries, which is what overwhelms a recovering service. The numbers behind attempt limits are discussed in setting retry budgets and max attempts.
Step 5 — Drop Work That No Longer Matters
After a long outage, some queued work is stale: a "your code expires in 5 minutes" email, a price-refresh job superseded by a newer one, a notification about something already resolved. Processing stale work wastes the recovering dependency's capacity. Give jobs a deadline and discard them past it.
@app.task(bind=True, acks_late=True)
def send_login_code(self, user_id, code_id, enqueued_at: float, ttl_s: int = 300):
if time.time() - enqueued_at > ttl_s:
STALE_DROPPED.labels(task="send_login_code").inc()
return "stale" # user has requested a new code by now
...
Deadlines turn the post-outage backlog into only the work that is still useful, which can be a large fraction for user-facing notifications. For superseded work, deduplicate by key so only the newest version runs — see deduplicating jobs with Redis SET NX keys.
Verification
Replay the incident in staging: block the dependency for 20 minutes under production-like load, unblock it, and record the dependency's request rate and error rate for the next 30 minutes.
# Request rate to the dependency must stay under its capacity through recovery
sum(rate(template_api_requests_total[1m]))
# Error rate should fall to baseline and stay there (no second outage)
sum(rate(template_api_requests_total{status=~"5.."}[1m])) / sum(rate(template_api_requests_total[1m]))
Pass criteria: no second outage, peak request rate below the configured cap, and the backlog drained within the time predicted by backlog / (cap − new arrivals) — the arithmetic in forecasting backlog drain time.
Gotchas & Edge Cases
Jitter in only one layer. Client libraries, job frameworks, and load balancers may each retry. Retries multiply across layers; disable or budget inner retries when the job framework retries.
Autoscalers reacting to backlog. Scaling on queue depth during an outage adds workers that can only wait. Cap maximum replicas at the dependency's capacity, or scale on processing rate rather than depth.
Clients with their own backoff. Some SDKs retry with their own synchronised backoff; configure them to fail fast so the job layer controls timing.
Retry timing stored at failure time. Jobs that computed their retry time during the outage keep it after recovery; changing the backoff policy mid-incident affects only new failures. That is another reason the concurrency cap, which applies at execution time, is the more dependable control during recovery.
Multiple queues sharing one dependency. Caps and ramps must be per dependency, not per queue, or three queues each ramping to "their" 40 slots add up to 120. Keep the semaphore or limiter keyed by dependency name.
Stale work that must not be dropped. Payments and orders are never "stale" — only drop work with explicit, business-approved deadlines.
FAQ
Is full jitter always better? For avoiding synchronised retries, yes. It increases the variance of an individual job's delay; if a maximum delay matters, cap the ceiling rather than removing jitter.
How do I pick the dependency's cap? From its documented limit, a load test, or its normal peak concurrency plus a margin. Start conservative and raise it while watching its latency.
What should the runbook say for the next outage? Three things: confirm the breaker is open and dependent consumption paused (so retries are not burning); do not manually scale workers up when the dependency recovers — the ramp handles it; and watch the dependency's error rate through the first ten minutes of recovery, ready to lower the cap if it degrades. Write it down before the incident, following writing runbooks for queue incidents.
Should retries go to a separate queue? It helps isolate retry traffic so new work is not delayed behind it, and makes retry volume visible. Combine with a lower concurrency cap on the retry queue.
Related
- Retry Strategies & Backoff — backoff shapes and attempt limits.
- Exponential Backoff with Jitter in Celery — full jitter in configuration.
- Circuit Breakers for Worker Dependencies — deciding when recovery starts.
- Testing Retry and Backoff Logic Deterministically — simulating the fleet before an outage.