Choosing a Go Job Queue Library
Picking a job queue for a Go service is mostly a question of where your data lives and what failure you can least afford, and this guide turns that into a repeatable decision as part of Task Queues in Go in Backend Frameworks & Worker Scaling. Rather than a feature checklist, it walks through the requirements that actually separate the options, scores them, and ends with a decision tree you can apply to the next service too.
Problem Statement
A platform team is standardising background processing across eight Go services. Today there are four approaches in use: Machinery on Redis in an older service, raw SQS polling in two, a hand-rolled Postgres queue in one, and goroutines with no durability in the rest. On-call engineers must learn every variant, and two incidents in the last quarter traced back to the goroutine-only services losing work during deploys. The team wants one default choice (with a documented exception path), chosen against explicit requirements rather than familiarity.
Prerequisites
- A list of the services and, for each, typical jobs per second, peak burst, longest job, and whether jobs are triggered by database writes.
- Knowledge of the infrastructure the team already operates well — Redis, Postgres, NATS, or a cloud provider's queue.
- Agreement on non-negotiables: for most teams, durability across deploys and at-least-once execution.
Step 1 — Write Down the Requirements That Decide
Most libraries tick the same boxes (retries, scheduling, a UI). The requirements below are the ones where they genuinely differ. Score each service against them.
# requirements.yaml — one block per service
service: invoicing
jobs_triggered_by_db_writes: always # always | sometimes | never
peak_jobs_per_second: 60
p99_pickup_latency_needed: 1s # how fast a job must start after enqueue
longest_job: 45s
infra_operated_well: [postgres] # what the team runs confidently today
needs_cross_service_fanout: false # one event consumed by several services?
needs_strict_ordering: per_invoice
jobs_triggered_by_db_writes is the most important line. When jobs are the consequence of a database write, the dual-write problem — commit succeeds, enqueue fails, or the reverse — is a correctness issue, and only a queue in the same database (or an outbox) solves it. needs_cross_service_fanout is the second: if several services consume the same event, you want a log or pub/sub system, not a job queue.
Step 2 — Map the Candidates
| Option | Store | Transactional enqueue | Throughput | Pickup latency | Ordering | Fan-out to many services |
|---|---|---|---|---|---|---|
| River | Postgres | Yes (InsertTx) |
Thousands/s | ~5–20 ms | Per-queue priority, no per-key | No |
| Asynq | Redis | No (outbox needed) | Tens of thousands/s | ~1 ms | Weighted queues, groups | No |
| NATS JetStream (pull consumer) | NATS | No | Very high | ~1 ms | Per-subject | Yes |
| SQS + Go poller | AWS | No | Scales with shards/calls | 10–20 ms | FIFO per group | Via SNS |
| Machinery | Redis/AMQP | No | Moderate | Varies | Limited | No |
| Goroutines + channel | Memory | No | Highest | Immediate | None | No |
The last row is included deliberately: in-process channels are the right answer for best-effort work (metrics batching, cache warming) and the wrong answer for everything a user depends on. Machinery works but is in maintenance mode, which is itself a reason not to standardise on it. NATS JetStream and SQS are message transports rather than job frameworks — you build retries and handlers yourself, as covered in NATS JetStream as a task queue.
Step 3 — Score Each Service
Weight the requirements by how much a violation would hurt, then score each option per service. A simple worksheet in code keeps it honest and reviewable:
// score.go — run with: go run score.go < services.json
type Option struct {
Name string
Transactional bool
MaxJobsPerSec int
PickupLatencyMs int
FanOut bool
InfraNeeded string
}
var options = []Option{
{"river", true, 3000, 20, false, "postgres"},
{"asynq", false, 30000, 2, false, "redis"},
{"jetstream", false, 100000, 2, true, "nats"},
{"sqs", false, 50000, 20, true, "aws"},
}
func score(s Service, o Option) (int, []string) {
total, notes := 0, []string{}
if s.JobsTriggeredByDBWrites == "always" {
if o.Transactional { total += 40 } else { notes = append(notes, "needs outbox") }
}
if s.PeakJobsPerSecond*3 <= o.MaxJobsPerSec { total += 20 } else { notes = append(notes, "throughput risk") }
if o.PickupLatencyMs <= s.PickupLatencyNeededMs { total += 10 }
if s.NeedsFanOut == o.FanOut || !s.NeedsFanOut { total += 15 }
if contains(s.InfraOperatedWell, o.InfraNeeded) { total += 15 } else { notes = append(notes, "new infra: "+o.InfraNeeded) }
return total, notes
}
The weights above encode one team's priorities — correctness of enqueue first, headroom second, operational familiarity third. Adjust them, but write them down; an explicit weighting ends the "which one do you like" debate quickly. A 3× headroom multiplier on peak throughput leaves room for growth and for replaying a backlog after an outage.
Step 4 — Read the Results as a Default Plus Exceptions
For the eight services in this scenario, the scores typically fall into two groups:
- Six services create jobs as a direct consequence of Postgres writes at well under a thousand jobs per second. River wins decisively: it removes the dual-write bugs and adds no infrastructure.
- Two services process high-volume event streams (clickstream enrichment at ~8,000 events per second) that are not tied to a database write and are consumed by several downstream teams. JetStream or SQS/SNS fit; a job queue does not.
That becomes the standard: River by default; a streaming transport when events fan out to several consumers or exceed a few thousand per second; Asynq when a service needs Redis-class latency and throughput for jobs not tied to DB writes. Goroutines only for explicitly best-effort work.
Step 5 — Prototype the Riskiest Assumption
Before standardising, test the assumption most likely to be wrong. For River on a shared primary, that is database impact at peak; for Asynq, it is Redis durability during failover. A short load test answers the first:
# Insert and work 100k no-op jobs at the target rate while watching the primary
go run ./cmd/loadgen --rate 300 --duration 10m --kind noop &
psql -c "SELECT relname, n_tup_ins, n_tup_upd, n_tup_del, n_dead_tup
FROM pg_stat_user_tables WHERE relname = 'river_job';"
# Also watch: replication lag, checkpoint frequency, p99 of the app's own queries
Run the test in three stages — expected peak, twice peak, three times peak — and record the application's own latency at each, not just the queue's. The number that decides the question is how much the application slows down when the queue is busy, because that is the cost the rest of the product pays for choosing a database-backed queue.
If application p99 latency moves by more than a few percent at three times expected peak, give the queue its own database or reconsider the choice for that service. The method is expanded in load testing queue throughput.
Step 6 — Write the Standard, Including the Migration Path
A standard that does not say how to get there from each current setup will be ignored. Record the default, the exception criteria, and one migration recipe per current approach:
## Background jobs in Go — standard (v1)
- Default: River, in the service's own Postgres database.
- Exception A (fan-out or > 3k jobs/s not tied to DB writes): NATS JetStream pull consumers.
- Exception B (sub-5ms pickup at high rate): Asynq with an outbox for DB-triggered jobs.
- Never for user-visible work: bare goroutines.
- Migration: dual-run per job kind behind a flag; old backend drains naturally.
- Shutdown: follow the drain/cancel pattern; grace period ≥ p99 job + 15s.
The dual-run migration is the same technique described in migrating from Redis to a Postgres job queue; the shutdown rule comes from graceful shutdown for Go workers.
Verification
A standard is working when incidents stop recurring and on-call load drops. Track for a quarter:
# Jobs lost or duplicated during deploys should be zero after migration
sum(increase(jobs_duplicate_side_effects_total[30d]))
sum(increase(jobs_enqueue_failures_total[30d]))
# Share of services on the standard backend
count(count by (service) (river_job_count)) / count(count by (service) (up{job="go-services"}))
Also review: how many runbooks exist for job systems (should shrink to one or two), and how long it takes a new engineer to debug a stuck job in any service.
Gotchas & Edge Cases
Standardising on the throughput leader by default. Picking the fastest option for every service imports the dual-write problem into the six services that never needed the throughput. Match the default to the common case.
Ignoring operational familiarity. A technically ideal choice that nobody on the team can operate at 3 a.m. is a worse choice. Weight it explicitly, as in Step 3.
Treating a transport as a framework. Choosing JetStream or SQS means building retries, dead-lettering, and handler routing. Budget for that work or wrap it in a small shared library.
Library longevity. Check release cadence and maintainer activity. Standardising on an unmaintained library turns a future security fix into an emergency migration.
FAQ
Is it bad to run more than one job system? Not inherently — one default plus a documented exception for streaming workloads is healthy. Four ad hoc systems with no criteria is the problem.
What about Temporal for Go? Temporal's Go SDK is excellent for long-running, multi-step workflows. It complements a job queue rather than replacing it; see when to move from job queues to Temporal.
Should the choice depend on team size? Indirectly. Smaller teams benefit more from fewer moving parts, which usually favours a database-backed queue in the database they already run.
Related
- Task Queues in Go — overview of the Go options.
- River: a Postgres Job Queue for Go — setting up the likely default.
- Getting Started with Asynq in Go — the Redis-backed alternative.
- Message Broker Comparison — broker-level trade-offs behind these libraries.