Sidekiq Quiet and Shutdown Timeouts

Sidekiq's shutdown is a two-phase sequence — stop fetching, then give running jobs a bounded time to finish — and deploys go wrong when that sequence and the orchestrator's timing disagree. This guide aligns them, as part of Graceful Shutdown & Worker Deployments in Backend Frameworks & Worker Scaling.

Problem Statement

A Rails app runs Sidekiq on Kubernetes with the default 30-second termination grace period and Sidekiq's default 25-second shutdown timeout. After each deploy, support sees a handful of customers with duplicated exports and a few "stuck" reports that never finished. Logs show Sidekiq pushing unfinished jobs back to Redis at shutdown, and some pods being SIGKILLed before that push completed. Long export jobs (up to 10 minutes) never finish during a deploy and restart from scratch every time. You want deploys that finish short jobs, hand back long ones promptly and cleanly, never lose a job to SIGKILL, and avoid restarting long work from zero.

Prerequisites

  • Sidekiq 7.x (OSS, Pro, or Enterprise — behaviour differs in Step 3).
  • Access to the Kubernetes Deployment spec (terminationGracePeriodSeconds, preStop).
  • Job duration percentiles per queue.
  • Idempotent jobs, since interrupted jobs run again.

Step 1 — Understand the Signals

Sidekiq responds to two signals during shutdown:

  • TSTP (quiet): stop fetching new jobs; keep running current ones. The process stays up.
  • TERM (shutdown): stop fetching, wait up to the shutdown timeout (-t, default 25 seconds) for running jobs, then push any still-running jobs back to Redis and exit.
# config/sidekiq.yml
:concurrency: 10
:timeout: 25          # the -t shutdown timeout: seconds to wait for running jobs on TERM
:queues:
  - critical
  - default
  - exports

The shutdown timeout must be shorter than the orchestrator's grace period, with room for Sidekiq to push unfinished jobs back and exit; otherwise SIGKILL arrives during the push.

TERM, wait, push back, exit On TERM, Sidekiq stops fetching immediately. Running jobs get up to the shutdown timeout of 25 seconds to finish. Jobs still running after that are pushed back to Redis and the process exits. The Kubernetes grace period must exceed the timeout plus the push-back time, or SIGKILL interrupts the push and jobs can be lost or duplicated. Sidekiq shutdown inside the grace period TERM running jobs finish (timeout 25 s) push back exit SIGKILL at grace end Rule: grace period ≥ timeout + push-back + margin (for example 25 s timeout, 45 s grace).

Step 2 — Align the Grace Period with the Timeout

Set the Kubernetes grace period well above Sidekiq's timeout. Add a short preStop so load balancers and autoscalers settle, and remember that preStop time counts against the grace period.

spec:
  terminationGracePeriodSeconds: 45      # > timeout (25) + preStop (5) + push-back margin
  containers:
    - name: sidekiq
      command: ["bundle", "exec", "sidekiq", "-C", "config/sidekiq.yml"]
      lifecycle:
        preStop:
          exec:
            command: ["/bin/sh", "-c", "sleep 5"]

Make sure Sidekiq is PID 1 (or started with exec from an entrypoint script), or the TERM signal may never reach it and every pod is SIGKILLed at the end of the grace period. The general drain pattern is described in Graceful Shutdown & Worker Deployments.

Step 3 — Know What Happens to Unfinished Jobs

What "push back" means depends on the edition:

Edition Fetch mechanism Job still running at timeout Process SIGKILLed
Sidekiq OSS BRPOP (basic fetch) Pushed back to queue Job lost
Sidekiq Pro super_fetch Pushed back to queue Recovered from private working queue on restart
Enterprise super_fetch + more Same as Pro Same as Pro

With OSS basic fetch, a job popped by a process that is SIGKILLed exists nowhere — the same at-most-once behaviour described in Redis Streams vs Redis Lists for job queues. Pro's super_fetch keeps in-progress jobs in a per-process working list and recovers them after a crash. If you run OSS, getting the timeout and grace period right is the only protection against losing jobs on deploy.

Where an in-flight job lives With OSS basic fetch, a job popped from the queue exists only in the worker's memory; if the process is SIGKILLed, the job is gone. With Pro super_fetch, the job is moved atomically into a private working list for that process; after a SIGKILL, the next Sidekiq process to start recovers jobs from orphaned working lists and re-queues them. Pod SIGKILLed with a job in flight OSS basic fetch job only in process memory lost, no error anywhere Pro super_fetch job in private working list recovered on next start Either way, avoiding SIGKILL through correct timing is better than recovering from it.
# Sidekiq Pro: enable reliable fetch
Sidekiq.configure_server do |config|
  config.super_fetch!
end

Step 4 — Quiet Early for Long-Running Queues

A 25-second timeout cannot accommodate 10-minute exports. Instead of stretching the timeout (and every deploy), quiet the export process long before stopping it, so it stops taking new exports and finishes the ones it has.

# Separate deployment for the exports queue, with a long grace period and early quiet
spec:
  terminationGracePeriodSeconds: 660     # exports can take 10 min
  containers:
    - name: sidekiq-exports
      command: ["bundle", "exec", "sidekiq", "-q", "exports", "-c", "3", "-t", "600"]
      lifecycle:
        preStop:
          exec:
            command: ["/bin/sh", "-c", "kill -TSTP 1 && sleep 20"]   # quiet, then allow TERM

Isolating long jobs in their own deployment means only that deployment rolls slowly; the main workers still deploy in seconds. For very long work, checkpointing (Step 5) is better than any timeout.

Separate long jobs from fast ones The main Sidekiq deployment handles critical and default queues with a 25-second timeout and a 45-second grace period, so it rolls quickly. The exports deployment receives TSTP in its preStop hook, stops taking new exports, and has an 11-minute grace period so running exports can finish before TERM ends them. Deploy behaviour by deployment main (critical, default) timeout 25 s, rolls in seconds exports TSTP running exports finish, no new ones fetched (up to 10 min) Only the export deployment rolls slowly; everything else deploys at normal speed.

Step 5 — Checkpoint Long Jobs So Interruption Is Cheap

Even with early quiet, node failures and emergency deploys interrupt long jobs. Make them resumable: record progress and skip completed work on retry.

class ExportJob
  include Sidekiq::Job
  sidekiq_options queue: "exports", retry: 5

  def perform(export_id)
    export = Export.find(export_id)
    batches = export.row_batches(size: 5_000)
    batches.each_with_index do |batch, i|
      next if i < export.completed_batches                 # resume point
      export.append_csv(batch)                             # idempotent per batch index
      export.update_column(:completed_batches, i + 1)
      return if Sidekiq::CLI.instance&.stopping?           # stop between batches on shutdown
    end
    export.finalize!
  end
end

Checking stopping? between batches lets the job return promptly when shutdown begins; because it returns without error, re-enqueue it explicitly (ExportJob.perform_async(export_id)) before returning, or let the push-back handle it if the timeout is reached. Either way, the next run resumes at the recorded batch instead of starting over.

Step 6 — Watch Deploys for Lost or Repeated Jobs

Measure what shutdown does in practice: jobs pushed back at shutdown, jobs retried after a deploy, and pods that exit with SIGKILL.

Sidekiq.configure_server do |config|
  config.on(:quiet)    { Rails.logger.info(event: "sidekiq.quiet") }
  config.on(:shutdown) { Rails.logger.info(event: "sidekiq.shutdown", busy: Sidekiq::WorkSet.new.size) }
end
# Pods killed rather than exiting cleanly (exit code 137)
sum(increase(kube_pod_container_status_last_terminated_exitcode{container=~"sidekiq.*"}[1d]) == 137)

Any SIGKILL of a Sidekiq OSS pod is a potential lost job; treat it as a bug in the timing configuration.

Step 7 — Keep Rollouts from Draining Capacity

During a rolling deploy, quieting pods still count as running but take no new work. If the rollout replaces many pods at once, effective capacity drops sharply while old pods drain. Keep maxUnavailable low and add surge capacity so new pods start before old ones quiet.

strategy:
  type: RollingUpdate
  rollingUpdate:
    maxSurge: 25%          # start replacements first
    maxUnavailable: 0      # never remove capacity before its replacement is ready

With readiness gated on Sidekiq having started (a simple probe that checks the process heartbeat in Redis), the rollout proceeds only as fast as new pods come up, and queue latency stays flat throughout the deploy. The broader rollout mechanics are in zero-downtime worker deploys on Kubernetes.

Verification

Run a deploy under load in staging with a mix of 1-second and 5-minute jobs:

kubectl rollout restart deployment/sidekiq
kubectl rollout restart deployment/sidekiq-exports
kubectl get pods -l app=sidekiq -w     # all old pods should terminate with exit code 0

Then reconcile: every job enqueued during the test completed exactly once (idempotency checks record duplicates), and every export finished without restarting from batch zero.

Gotchas & Edge Cases

Shell entrypoints swallowing TERM. sh -c "bundle exec sidekiq" makes the shell PID 1; it does not forward signals. Use exec or run Sidekiq directly.

Timeout larger than grace. A -t 60 with a 30-second grace period guarantees SIGKILL mid-shutdown.

Autoscaler scale-down. Scale-down evicts pods the same way as deploys; long jobs need the same protection, or scale-down stabilization windows long enough to avoid churn.

Scheduled and retry sets are safe. Jobs waiting in the scheduled or retry sets live in Redis, not in the process, and are unaffected by shutdown.

FAQ

Is TSTP necessary if TERM already stops fetching? For short jobs, TERM alone is fine. TSTP is useful when you want a long quiet period before the stop — for long-running queues — or during manual maintenance.

Should I raise the timeout for all workers? No. It slows every deploy and hides long jobs. Split long jobs into their own deployment and checkpoint them.

How do I pick the timeout for the main deployment? Take the p99 duration of jobs on the queues that deployment serves and add a margin; jobs that exceed it are pushed back and rerun, which is acceptable if they are rare and idempotent. If the p99 is above a minute, the queue probably mixes long and short jobs and should be split, as in Step 4.

What happens to jobs pushed back at shutdown? They go to the front of their queue and are picked up by another process almost immediately, with no retry count consumed — the interrupted run simply did not happen from Sidekiq's point of view. Any side effects it performed before interruption did happen, which is why idempotency matters.

Can Kubernetes tell Sidekiq to quiet on its own? Only through the preStop hook, as shown in Step 4. There is no built-in signal for "stop taking work but keep running"; TSTP from preStop is the standard way to get that phase.

Does super_fetch make timing irrelevant? It prevents loss after SIGKILL, but interrupted jobs still restart. Timing still decides how often that happens.

Related