Sidekiq Queue Weights and Capsules

A Sidekiq process listens to several queues and has to decide, every time a thread is free, which queue to take the next job from. Queue weights shape that decision; capsules, added in Sidekiq 7, go further and give queues their own thread pools inside one process. This guide shows how both work and when each is the right tool, as part of Sidekiq Performance Tuning in Backend Frameworks & Worker Scaling.

Problem Statement

A Rails application runs Sidekiq with concurrency: 10 and three queues: critical (password resets, payment webhooks), default, and low (report generation, imports). The configuration lists them in strict order. On quiet days everything is fine. During a marketing campaign, critical receives a steady stream of webhook jobs, and low stops being processed for six hours — reports requested in the morning arrive in the evening. The team switched to weights, and then the opposite happened: a large import in low occupied threads for 20 minutes at a time, and payment webhooks waited. They also have a job that calls a legacy API which tolerates only one request at a time. They want urgent work to stay fast, low-priority work to keep moving, and a way to cap concurrency for specific job types.

Prerequisites

  • Sidekiq 7.0 or newer (capsules require 7).
  • Queue latency per queue (Sidekiq::Queue.new("critical").latency) in your monitoring.
  • Typical and worst-case job durations per queue.
  • A Redis connection pool large enough for the total concurrency you plan — see tuning the Sidekiq Redis connection pool.

Step 1 — Understand Strict Ordering vs Weights

When queues are listed without weights, Sidekiq checks them in strict order: it takes from critical if anything is there, otherwise default, otherwise low. As long as critical is never empty, low is never served.

# strict: starves lower queues under sustained load
:queues:
  - critical
  - default
  - low

With weights, Sidekiq builds a list in which each queue appears as many times as its weight, shuffles it on every fetch, and checks queues in that shuffled order. A queue with weight 6 among weights 6, 3, 1 is checked first 60% of the time — but when it is empty, the next queue in the shuffle is tried, so no capacity is wasted.

# weighted: every queue makes progress
:queues:
  - [critical, 6]
  - [default, 3]
  - [low, 1]
Share of fetches under sustained load When all three queues have work waiting, strict ordering sends every fetch to critical, so default and low get nothing. With weights of 6, 3 and 1, critical is checked first 60 percent of the time, default 30 percent and low 10 percent, so every queue keeps moving. When a queue is empty its share goes to the others. Who gets the next free thread when every queue is busy strict order critical 100% default 0%, low 0% — starved weights 6/3/1 critical 60% default 30% low low still gets about 10% of fetches Shares are of fetches, not of thread time: a long job holds its thread far longer.

Weights distribute fetches, not thread time. A low job that runs for 20 minutes holds its thread for all of that time, while a critical job that takes 50 ms frees its thread almost at once. Over a few minutes, long jobs accumulate: with 10 threads and 10% of fetches going to 20-minute imports, most threads end up busy with imports. That is exactly what happened to the payment webhooks.

Step 2 — Choose Weights from Latency Targets

Start from what each queue needs, not from a feeling about importance:

  • critical: latency target under a few seconds; short jobs.
  • default: latency target of a minute; mixed durations.
  • low: latency target of an hour; long jobs.

Give short, urgent queues high weights, and keep long-running jobs out of shared processes entirely (Step 3). Ratios like 6/3/1 or 10/5/1 are common. After changing weights, watch latency per queue for a week: if critical latency rises during bursts, the problem is usually total capacity or long jobs, not the weight. Weighting a queue heavily when it is empty costs nothing, so err towards higher weights for urgent queues.

Step 3 — Isolate Queues with Capsules

Sidekiq 7 capsules divide one process into independent thread pools, each with its own queues and concurrency. Long-running work gets a fixed number of threads it can never exceed, and urgent work keeps its own threads regardless of what else is running:

# config/initializers/sidekiq.rb
Sidekiq.configure_server do |config|
  # the default capsule: urgent and normal work
  config.queues = %w[critical,6 default,3]
  config.concurrency = 10

  # long jobs cannot take more than 3 threads
  config.capsule("bulk") do |cap|
    cap.concurrency = 3
    cap.queues = %w[low imports]
  end

  # the legacy API accepts one request at a time
  config.capsule("legacy-serial") do |cap|
    cap.concurrency = 1
    cap.queues = %w[legacy_api]
  end
end
Capsules inside one Sidekiq process A single Sidekiq process contains three capsules. The default capsule has 10 threads serving the critical and default queues with weights 6 and 3. The bulk capsule has 3 threads serving low and imports, so long jobs can never occupy more than three threads. The legacy-serial capsule has a single thread for the legacy API queue, so its jobs run one at a time. All capsules share the process memory and the Redis connection budget. One process, three independent thread pools sidekiq process (14 threads total) default capsule 10 threads critical ×6, default ×3 never blocked by long jobs bulk 3 threads low, imports legacy-serial 1 thread legacy_api Capsules share memory and the GIL; they isolate thread slots, not CPU.

A capsule with concurrency = 1 gives serial execution per process. With several processes, you get one concurrent job per process, not one globally; for a strict global limit, run the capsule in exactly one process or use a distributed limiter such as Sidekiq Enterprise's Sidekiq::Limiter.concurrent.

Step 4 — Know When to Split into Separate Processes

Capsules isolate thread slots, but all threads in a Ruby process share one GVL (global VM lock) and one heap. A capsule running CPU-heavy report generation still slows every other capsule in the process, and a job that bloats memory affects them all. Split into separate processes (or Kubernetes deployments) when:

  • A queue's jobs are CPU-bound or memory-hungry — they need their own resource limits.
  • Queues need to scale independently, with their own autoscaling signal.
  • A queue needs a different deploy or restart cadence (for example, imports that must not be interrupted by every web deploy).
# separate deployments, each with its own concurrency and replicas
bundle exec sidekiq -C config/sidekiq_urgent.yml    # critical, default
bundle exec sidekiq -C config/sidekiq_bulk.yml      # low, imports

Capsules are ideal for lightweight isolation and concurrency caps without extra infrastructure; separate processes are the stronger tool when resources or scaling differ. Many teams use both — separate deployments for bulk work, and a serial capsule inside the urgent deployment for the legacy API. Diagnosing Sidekiq memory bloat helps decide whether a queue belongs in its own process.

Step 5 — Size Redis Connections for Capsules

Each capsule has its own Redis connection pool sized from its concurrency. The total connections per process is roughly the sum of the capsules' concurrency plus a few for the process's own housekeeping. Multiply by the number of processes and check it against Redis maxclients and any provider connection limit before rolling out. Adding a capsule with 3 threads to 40 processes adds around 120 connections.

Step 6 — Monitor Latency per Queue and Utilisation per Capsule

The only reliable proof that weights and capsules are right is latency per queue against its target:

Sidekiq::Queue.all.each do |q|
  StatsD.gauge("sidekiq.queue.latency", q.latency, tags: ["queue:#{q.name}"])
  StatsD.gauge("sidekiq.queue.size", q.size, tags: ["queue:#{q.name}"])
end

Also watch busy threads per capsule (from Sidekiq::WorkSet, grouped by queue). A bulk capsule that is always 3/3 busy with growing latency needs more threads or more processes; a default capsule that is rarely more than half busy can give threads back. More on the metrics in exporting Sidekiq metrics to Prometheus.

Queue latency before and after capsules Before the change, strict ordering and then weights alone produced a p95 latency of 4 minutes for critical during imports and 6 hours for low during campaigns. After moving long jobs into a bulk capsule with 3 threads and weighting critical and default 6 to 3, critical p95 latency is 2 seconds and low is 25 minutes, both within target. p95 queue latency: before vs after before after target critical 4 min 2 s < 5 s low 6 h 25 min < 1 h Measured during a campaign with an import running.

Verification

  • With every queue loaded, each queue's latency stays within its target; no queue's latency grows without bound.
  • During a large import, critical p95 latency does not change.
  • The legacy API sees at most one concurrent request per process (or globally, if configured).
  • Redis connected clients stay below 80% of maxclients at full scale.

Gotchas & Edge Cases

Equal weights mean random, not round-robin. [a, 1], [b, 1] shuffles on every fetch; it does not alternate strictly. Over many fetches the split is even.

A queue listed in two capsules. Sidekiq allows it, and both capsules then fetch from it. That can be useful for overflow, but it defeats a concurrency cap. Keep capped queues in exactly one capsule.

Weights in sidekiq.yml vs code. Configuration in the initializer overrides the YAML for the default capsule. Keep one source of truth to avoid a deploy that silently reverts to strict ordering.

Shutdown timeouts apply per process, not per capsule. A bulk capsule running a 20-minute import delays the shutdown of the whole process, including the urgent capsule, during every deploy. If long jobs cannot be made resumable, that is another reason to move them to their own deployment.

Reliable fetch. Sidekiq Pro's super_fetch works per capsule; make sure it is enabled for every capsule, or jobs in one capsule can be lost on a crash while others are protected.

FAQ

Are capsules available in open-source Sidekiq? Yes, capsules are part of Sidekiq 7 itself. Global concurrency limiters and super_fetch are Pro and Enterprise features.

Should every queue get its own capsule? No. Capsules add Redis connections and fragment thread capacity, so idle threads in one capsule cannot help a busy one. Use them for work that must be capped or protected, and let weights handle the rest.

How is this different from priority within a queue? Weights and capsules choose between queues. Sidekiq does not order jobs within a queue by priority; for that, route urgent jobs to a separate queue, as in Sidekiq middleware for job prioritization.

Related