Celery Task Time Limits

A task with no time limit can hold a worker slot forever — stuck on a socket without a timeout, a lock, or an unexpectedly huge input. Celery's soft and hard time limits bound every task, and this guide sets them up so that stuck tasks are stopped, cleaned up, and retried or reported, as part of Celery Architecture & Configuration in Backend Frameworks & Worker Scaling.

Problem Statement

A document-conversion service runs 12 Celery workers with 8 prefork processes each. Every few days, throughput collapses: dozens of processes are busy for hours on conversions that normally take 20 seconds, stuck on a third-party converter that stopped responding without closing the connection. The queue backs up until someone restarts workers by hand. When a restart kills a task mid-write, partial output files are left behind. You want every task bounded in time, a chance to clean up before being stopped, a hard stop for tasks that ignore the first signal, limits appropriate to each task type, and stuck tasks visible in metrics.

Prerequisites

  • Celery 5.3+ with the prefork pool (hard time limits require it; see Step 6 for other pools).
  • Measured duration percentiles per task type.
  • Task code that can clean up partial work (temporary files, database transactions).
  • task_acks_late configured as in Celery acks_late and worker crash safety.

Step 1 — Understand Soft and Hard Limits

Celery has two limits:

  • Soft time limit: when reached, Celery raises SoftTimeLimitExceeded inside the task. The task can catch it, clean up, and exit — or re-raise to fail.
  • Hard time limit: when reached, the pool process running the task is killed and replaced. No cleanup code runs.
# celeryconfig.py — global defaults
task_soft_time_limit = 300      # 5 min: task gets SoftTimeLimitExceeded
task_time_limit = 330           # 5.5 min: process killed if still running

Set the hard limit a little above the soft limit, leaving time for cleanup. The soft limit is the one you design around; the hard limit is the backstop for code that swallows the exception or is stuck in native code where the exception cannot be delivered.

Two limits, two behaviours A normal task finishes in twenty seconds. A stuck task reaches the soft limit at five minutes, when SoftTimeLimitExceeded is raised inside it so it can delete temporary files and fail cleanly. If it ignores the exception or is stuck in native code, the hard limit at five and a half minutes kills the pool process, and a new process takes its place. soft 300 s, hard 330 s normal 20 s, done stuck waiting on converter soft hard kill Soft: exception inside the task, cleanup runs. Hard: process replaced, no cleanup.

Step 2 — Catch the Soft Limit and Clean Up

Handle SoftTimeLimitExceeded where you own resources, clean them up, and let the task fail or retry.

from celery.exceptions import SoftTimeLimitExceeded

@app.task(bind=True, acks_late=True, soft_time_limit=120, time_limit=150,
          autoretry_for=(ConverterTimeout,), max_retries=3, retry_backoff=True)
def convert_document(self, doc_id: int):
    tmp = tempfile.mkdtemp(prefix=f"conv-{doc_id}-")
    try:
        src = download(doc_id, tmp)
        out = converter.convert(src, timeout=90)          # client timeout below the soft limit
        upload_result(doc_id, out)
    except SoftTimeLimitExceeded:
        log.warning("conversion exceeded soft limit", doc_id=doc_id, attempt=self.request.retries)
        raise self.retry(countdown=60, exc=ConverterTimeout("soft limit"))
    finally:
        shutil.rmtree(tmp, ignore_errors=True)             # runs on success, error, and soft limit

Two layers are visible here: a client-side timeout (90 s) that should normally fire first, and the soft limit (120 s) that catches anything the client timeout missed. Time limits are a safety net, not a substitute for timeouts on every network call — the stuck-socket problem in the scenario is best fixed at the client.

Step 3 — Set Limits per Task Type

One global limit either kills legitimately long tasks or lets short ones hang for too long. Set a conservative global default and override per task from measured durations — roughly 3–5× the p99, with an absolute cap.

# per-task via decorator
@app.task(soft_time_limit=30, time_limit=40)
def send_email(...): ...

@app.task(soft_time_limit=3600, time_limit=3660)
def build_monthly_report(...): ...

# or centrally via annotations, keeping limits in one reviewed place
task_annotations = {
    "mail.send_email":             {"soft_time_limit": 30,   "time_limit": 40},
    "docs.convert_document":       {"soft_time_limit": 120,  "time_limit": 150},
    "reports.build_monthly_report": {"soft_time_limit": 3600, "time_limit": 3660},
}
# p99 duration per task, as input for limits
histogram_quantile(0.99, sum by (le, task) (rate(celery_task_runtime_seconds_bucket[7d])))

Recheck limits when task behaviour changes; a limit set from last year's data can start killing healthy tasks as inputs grow. Duration histograms are covered in choosing histogram buckets for job duration.

Step 4 — Understand What a Hard Kill Leaves Behind

A hard kill terminates the pool process immediately. Open transactions roll back (the database connection closes), but anything outside a transaction — files written, external API calls made, messages published — stays as it was. With late ack and task_reject_on_worker_lost, the task is requeued and runs again.

# Write output atomically so a hard kill never leaves a half-written artifact visible
def upload_result(doc_id, local_path):
    tmp_key = f"converted/{doc_id}.pdf.part-{uuid4().hex}"
    storage.put(tmp_key, local_path)
    storage.copy(tmp_key, f"converted/{doc_id}.pdf")       # visible only when complete
    storage.delete(tmp_key)
# A lifecycle rule deletes stray *.part-* objects after 1 day
What a hard kill leaves behind When the pool process is killed, the open database transaction rolls back because its connection closes. Files already written and external API calls already made persist. Writing output to a temporary part key and copying it to the final key only when complete means a killed task leaves a stray part file for the lifecycle rule to clean up, never a half-written final artifact. After a hard kill DB transaction rolled back: safe files, API calls persist: need idempotency .part + copy final file all-or-nothing A lifecycle rule deletes stray .part objects; the retried task writes a fresh one.

Design for the hard kill as if it were a crash, because it is one: atomic final writes, idempotent retries, and periodic cleanup of temporary artifacts.

Step 5 — Keep Limits Consistent with Broker Timeouts

Time limits interact with the broker's redelivery timers:

  • Redis broker: visibility_timeout must exceed the longest time_limit (plus any countdown), or a still-running task is redelivered to a second worker before its own limit stops it.
  • RabbitMQ: consumer_timeout (default 30 minutes) closes the channel for deliveries held longer than that; tasks with longer limits need a queue policy raising it, as covered in RabbitMQ consumer timeout for unacked messages.
LONGEST_TIME_LIMIT = 3660
broker_transport_options = {"visibility_timeout": LONGEST_TIME_LIMIT + 600}   # Redis

A simple startup assertion in the worker that compares the largest configured time_limit with the broker timeout catches drift when someone raises a limit.

Timers must nest The converter client timeout of 90 seconds sits inside the soft limit of 120 seconds, which sits inside the hard limit of 150 seconds, which must sit inside the broker's redelivery timer: Redis visibility timeout or RabbitMQ consumer timeout. If the broker timer is shorter than the hard limit, a still-running task is redelivered to a second worker. client < soft < hard < broker broker visibility / consumer timeout hard limit 150 s soft limit 120 s client timeout 90 s The innermost timer should fire in normal failures; outer ones catch what it misses.

Step 6 — Know Which Pools Support Which Limits

Hard time limits rely on killing a child process, so they work only with the prefork pool. With solo, threads, gevent, or eventlet, the hard limit cannot be enforced, and the soft limit's exception delivery varies.

Pool Soft limit Hard limit
prefork Yes (signal to child) Yes (child killed)
threads No No
gevent / eventlet Timeout raised in greenlet (cooperative) No
solo No No

For non-prefork pools, rely on client timeouts and on gevent.Timeout or explicit deadline checks in the task, and consider moving tasks that can hang in native code to a prefork worker. Pool choice is covered in choosing Celery prefork vs gevent pools.

Step 7 — Make Timeouts Visible

Count soft-limit and hard-limit events per task; a rising rate is an early signal of a degrading dependency or growing inputs.

from celery.signals import task_failure

@task_failure.connect
def on_failure(sender=None, exception=None, **_):
    if isinstance(exception, SoftTimeLimitExceeded):
        TIME_LIMIT_EVENTS.labels(task=sender.name, kind="soft").inc()
    elif isinstance(exception, TimeLimitExceeded):
        TIME_LIMIT_EVENTS.labels(task=sender.name, kind="hard").inc()
sum by (task, kind) (increase(celery_time_limit_events_total[1h])) > 5

Hard-limit events deserve more attention than soft ones: each means cleanup code did not run and a process was replaced.

Verification

def test_soft_limit_cleans_up(celery_worker, monkeypatch):
    monkeypatch.setattr(converter, "convert", lambda *a, **k: time.sleep(10))
    res = convert_document.apply_async(args=[1], soft_time_limit=1, time_limit=3)
    with pytest.raises(Exception):
        res.get(timeout=10)
    assert not glob.glob("/tmp/conv-1-*")                     # temp dir removed

In staging, point the converter at a black-hole endpoint and confirm that tasks stop at the soft limit, retry with backoff, and that worker throughput recovers without a manual restart.

Gotchas & Edge Cases

Catching broad exceptions. except Exception: catches SoftTimeLimitExceeded and may swallow it, leaving the task running until the hard kill. Re-raise it or handle it explicitly first.

Limits on the caller side. result.get(timeout=...) limits how long a caller waits, not how long the task runs. Both are needed.

Limits and worker_max_tasks_per_child. A hard kill replaces the process, which also resets any per-child memory growth. That is a side effect, not a strategy: use worker_max_tasks_per_child or worker_max_memory_per_child for leaks, as in fixing Celery worker memory leaks.

Retries reset the clock. Each retry gets a fresh limit. Bound total effort with max_retries.

time_limit below soft_time_limit. If the hard limit is lower, the soft limit never fires. Keep hard above soft.

FAQ

What limit should I start with? A global soft limit of a few minutes and hard limit slightly above, then per-task overrides from measured p99 durations. Anything genuinely longer than about an hour should be split into steps.

Can a task extend its own limit? No. Split long work into chained steps, each with its own limit, and checkpoint progress between them.

Why did my task hit the hard limit even though it catches the soft one? The soft limit is delivered as a signal-raised exception in the main thread of the pool process; if that thread is inside native code (a C extension, a blocking system call without a timeout), the exception is not raised until control returns to Python — which may be never. The hard limit then fires. Fix it by adding timeouts to the underlying calls so control returns regularly.

Does a hard kill count as a retry? With late ack and reject_on_worker_lost, the message is requeued and the task runs again as a redelivery, not through self.retry, so max_retries does not bound it. Track redeliveries separately.

Related