Cutting Worker Costs with Spot Instances
Queue workers are among the best workloads for spot and preemptible instances — stateless, horizontally scaled, and built around redelivery — and this guide shows how to move them there without turning interruptions into incidents, as part of Capacity Planning for Job Queues in Observability & Monitoring for Job Queues. Spot capacity is typically 60–90% cheaper than on-demand; the price is that the cloud provider can reclaim a node with about two minutes' notice (AWS, Azure) or 30 seconds (Google Cloud).
Problem Statement
A media company runs 120 worker nodes on on-demand instances for thumbnail generation, transcoding, and metadata extraction; compute is its largest infrastructure cost. Finance asks for a 40% reduction. The engineering concern is interruptions: transcodes run up to 20 minutes, and an early experiment that moved a transcoding pool to spot caused jobs to restart repeatedly, some never finishing during a period of heavy reclamation. You want most worker capacity on spot, interruptions handled so that short jobs finish and long jobs checkpoint, a guaranteed on-demand floor for critical queues, and a measured cost saving net of redelivered work.
Prerequisites
- Kubernetes with node groups or node pools that can use spot capacity (EKS managed node groups or Karpenter, GKE spot node pools, AKS spot pools).
- Workers that already shut down gracefully on
SIGTERMand whose jobs are idempotent — see graceful shutdown for Go workers or handling SIGTERM in Celery workers. - An interruption handler: AWS Node Termination Handler (or Karpenter's native handling), GKE's built-in graceful node shutdown, or Azure's scheduled events handling.
- Job duration percentiles per job type.
Step 1 — Classify Jobs by Interruption Tolerance
Not every job belongs on spot. The deciding factor is how much work an interruption throws away compared with the notice period.
| Job type | p99 duration | Idempotent | Checkpointable | Placement |
|---|---|---|---|---|
| Thumbnails | 4 s | Yes | Not needed | Spot |
| Metadata extraction | 30 s | Yes | Not needed | Spot |
| Transcode | 20 min | Yes | Yes (per segment) | Spot with checkpoints |
| Payment reconciliation | 2 min | Yes, with keys | No | On-demand |
| Customer-facing exports | 5 min | Yes | Partially | On-demand floor + spot burst |
Jobs that finish well within the notice period lose nothing: the worker stops fetching on the notice and completes what it has. Long jobs lose everything since their last checkpoint, so they need checkpoints to qualify. Jobs whose redelivery is expensive or visible (payments, anything with a strict latency objective) keep an on-demand floor.
Step 2 — Create Separate Spot and On-Demand Pools
Separate node pools let you place each queue's workers deliberately. With Karpenter on EKS, two NodePools express it cleanly; the spot pool allows many instance types so the provider has more capacity pools to draw from, which reduces interruption rates.
apiVersion: karpenter.sh/v1
kind: NodePool
metadata: { name: workers-spot }
spec:
template:
metadata: { labels: { capacity: spot } }
spec:
requirements:
- { key: karpenter.sh/capacity-type, operator: In, values: ["spot"] }
- { key: karpenter.k8s.aws/instance-category, operator: In, values: ["c", "m", "r"] }
- { key: karpenter.k8s.aws/instance-generation, operator: Gt, values: ["5"] }
- { key: kubernetes.io/arch, operator: In, values: ["amd64", "arm64"] } # widen the pool
taints: [{ key: capacity, value: spot, effect: NoSchedule }]
disruption: { consolidationPolicy: WhenEmptyOrUnderutilized }
---
apiVersion: karpenter.sh/v1
kind: NodePool
metadata: { name: workers-ondemand }
spec:
template:
metadata: { labels: { capacity: on-demand } }
spec:
requirements:
- { key: karpenter.sh/capacity-type, operator: In, values: ["on-demand"] }
limits: { cpu: "200" } # the floor is deliberate and bounded
Diversifying instance types and architectures matters more than any other spot setting: a pool restricted to one instance type in one zone is reclaimed together when that capacity pool tightens.
Step 3 — Place Workers with Tolerations and a Floor
Spot-tolerant workers get a toleration and a preference for spot; critical workers are pinned to on-demand. For queues that need both, run two deployments — a small on-demand floor and a spot deployment that scales with backlog.
# Thumbnail workers: spot only
spec:
template:
spec:
tolerations: [{ key: capacity, value: spot, effect: NoSchedule }]
nodeSelector: { capacity: spot }
terminationGracePeriodSeconds: 110 # < 120 s notice minus handler overhead
---
# Export workers: on-demand floor (2 replicas) ...
spec:
replicas: 2
template:
spec:
nodeSelector: { capacity: on-demand }
---
# ... plus a spot burst deployment scaled by KEDA on the same queue
spec:
template:
spec:
tolerations: [{ key: capacity, value: spot, effect: NoSchedule }]
nodeSelector: { capacity: spot }
The floor guarantees progress even during a wave of reclamation; the burst deployment provides cheap capacity most of the time. Scaling the burst deployment on queue depth is covered in scaling workers with KEDA on queue length.
Step 4 — Turn the Interruption Notice into a Graceful Drain
The notice reaches the node, not your process. An interruption handler (Karpenter or AWS Node Termination Handler) watches for it, cordons the node, and evicts pods — which sends SIGTERM to your worker with the pod's grace period. The worker must stop fetching immediately and finish or hand back in-flight jobs within the remaining time.
# celery worker settings for spot nodes
worker_prefetch_multiplier = 1 # don't hold messages you may not get to finish
task_acks_late = True # unacked work returns to the queue if we're killed
task_reject_on_worker_lost = True
worker_cancel_long_running_tasks_on_connection_loss = True
# In long tasks: check a shutdown flag between checkpoints
import signal, threading
shutting_down = threading.Event()
signal.signal(signal.SIGTERM, lambda *_: shutting_down.set())
@app.task(bind=True, acks_late=True)
def transcode(self, video_id: str):
for seg in pending_segments(video_id): # resumes from the last checkpoint
if shutting_down.is_set():
raise self.retry(countdown=5) # hand back promptly, keep progress
encode_segment(video_id, seg)
mark_segment_done(video_id, seg) # durable checkpoint
finalize(video_id)
Set terminationGracePeriodSeconds slightly below the notice period so the pod exits before the node disappears. On GKE, where the notice is about 30 seconds, only short jobs and aggressively checkpointed ones fit; plan the placement table accordingly.
Step 5 — Watch Interruption Rates and Redelivered Work
Savings are real only if interruptions do not waste much work. Track interruptions, redelivered jobs, and work thrown away, per queue.
# Interruptions per hour (from Karpenter or the termination handler)
sum(increase(karpenter_nodes_terminated_total{reason="interruption"}[1h]))
# Share of jobs that are redeliveries (retries caused by worker loss)
sum(rate(jobs_redelivered_total{cause="worker_lost"}[1h])) / sum(rate(jobs_completed_total[1h]))
# Wasted compute: seconds of work discarded by interruptions
sum(increase(job_discarded_work_seconds_total[1h]))
If redelivery rates for a queue exceed a few percent, the job type is too long or poorly checkpointed for the current interruption rate — move it to on-demand or add checkpoints. A sudden rise across all pools means the chosen instance types are under pressure; widen the diversification in Step 2.
Step 6 — Compute the Net Saving
Compare cost per thousand completed jobs before and after, including the redelivered work.
def cost_per_1k_jobs(node_hours: float, price_per_hour: float, jobs_completed: int) -> float:
return node_hours * price_per_hour / jobs_completed * 1000
before = cost_per_1k_jobs(node_hours=120 * 720, price_per_hour=0.34, jobs_completed=410_000_000)
after = cost_per_1k_jobs(node_hours=24 * 720, price_per_hour=0.34, jobs_completed=410_000_000) \
+ cost_per_1k_jobs(node_hours=104 * 720, price_per_hour=0.11, jobs_completed=410_000_000)
print(f"{before:.4f} -> {after:.4f} per 1k jobs, saving {(1 - after / before):.0%}")
# 0.0716 -> 0.0344 per 1k jobs, saving 52%
The spot pool needed slightly more nodes (104 spot plus a 24-node on-demand floor, instead of 120) to cover redelivered work and churn, and the saving still exceeded the 40% target. Revisit the calculation quarterly: spot prices and interruption rates change.
Verification
Before moving production queues, run an interruption drill in staging: use the provider's fault-injection tooling (AWS FIS aws:ec2:send-spot-instance-interruptions) or simply drain spot nodes with the same grace period, under load.
aws fis start-experiment --experiment-template-id "$SPOT_INTERRUPT_TEMPLATE"
# Then check: no job failed permanently, redeliveries only for in-flight jobs, transcodes resumed
Pass criteria: zero jobs lost, redeliveries limited to jobs in flight at the notice, long jobs resumed from checkpoints rather than restarting, and queue wait staying within its objective thanks to the on-demand floor.
Gotchas & Edge Cases
Prefetched messages. A worker holding prefetched messages releases them only when its connection closes or the visibility timeout expires. With prefetch_multiplier=1 and late acks, loss of a node costs at most one message per slot.
Zone concentration. Spot capacity can vanish across a whole zone. Spread spot pools across zones and keep the on-demand floor multi-zone too.
Daemon overhead. Every node runs log agents and exporters; many small spot nodes multiply that overhead. Prefer medium-sized instances.
Long grace periods. A pod with a 600-second grace period on a node reclaimed at 120 seconds simply dies. Keep grace periods within the notice window on spot pools.
FAQ
Are spot interruptions frequent? It varies by instance type, region, and time; diversified pools commonly see a few percent of nodes interrupted per day. Measure your own rate — it drives the placement table.
Should the broker or database run on spot? No. Stateful components with failover costs belong on on-demand or reserved capacity; only stateless workers should run on spot.
What about Savings Plans or reserved instances? Use them for the on-demand floor, which runs all the time. Spot covers the variable part above the floor.
Related
- Capacity Planning for Job Queues — sizing the floor and the burst.
- Scaling Workers to Zero — the other half of cost control.
- Zero-Downtime Worker Deploys on Kubernetes — the same drain mechanics used for deploys.
- Configuring Visibility Timeouts for Long-Running Workers — checkpoints and lease renewal for long jobs.