Cloud Pub/Sub Push vs Pull for Workers
Google Cloud Pub/Sub can deliver the same messages to your workers in two ways — pushing them to an HTTP endpoint or letting workers pull them — and the choice changes how you control concurrency, handle slow jobs, and scale. This guide compares the two on a concrete workload as part of Managed Cloud Queues for Background Jobs in Backend Frameworks & Worker Scaling, then configures each correctly: ack deadlines, flow control, dead-letter topics, retry policy, and exactly-once delivery.
Problem Statement
A logistics company publishes ShipmentScanned events to a Pub/Sub topic. Three teams consume them: a tracking service updates customer-visible status (fast, ~50 ms per event), a billing service computes surcharges (moderate, ~300 ms), and an analytics job enriches events with route data (slow, up to 20 seconds, calls an external mapping API). All three started with push subscriptions to Cloud Run because it needed no worker code. Tracking works perfectly. Billing occasionally double-charges. Analytics times out constantly, and its retries hammer the mapping API. Each team needs to know whether push or pull fits its workload, and how to configure it.
Prerequisites
- A Pub/Sub topic and permission to create subscriptions on it.
- For push: an HTTPS endpoint (Cloud Run, GKE, or App Engine) and a service account for Pub/Sub to authenticate as.
- For pull: a long-running worker (GKE, Compute Engine, or Cloud Run jobs/worker pools) using the
google-cloud-pubsubclient with streaming pull. - A dead-letter topic per subscription, and the Pub/Sub service agent granted publisher on it and subscriber on the source subscription.
Step 1 — Understand How Each Model Delivers
A subscription is the unit of delivery: each subscription gets its own copy of every message published to the topic, so the three teams each create their own. Within a subscription, messages are distributed to consumers, and each must be acknowledged within the ack deadline or it is redelivered.
Push: Pub/Sub sends an HTTPS POST per message. A 2xx response is an ack; anything else, or no response within the ack deadline, is a nack and triggers redelivery after the retry policy's backoff. Pub/Sub adjusts the push rate itself, slowing down when your endpoint returns errors ("push backoff") and speeding up when it succeeds.
Pull (streaming): your worker holds a long-lived gRPC stream; Pub/Sub sends messages down it up to your client's flow-control limits, and your code acks each one explicitly. The client library extends ack deadlines automatically while messages are being processed.
Step 2 — Match the Workload to the Model
The three teams' results follow directly from the delivery mechanics:
| Workload | Push fits? | Why |
|---|---|---|
| Tracking: 50 ms, stateless, high rate | Yes | Short handlers finish well inside the ack deadline; Cloud Run autoscaling tracks the push rate |
| Billing: 300 ms, must not double-charge | Either, with care | Duplicates came from non-idempotent code, not the model — fix with exactly-once delivery or idempotency (Step 5) |
| Analytics: up to 20 s, rate-limited external API | Pull | Needs a hard cap on in-flight work to respect the API's limit; push concurrency is governed by Pub/Sub and Cloud Run scaling, not by you |
The general rule: push when handlers are short and you are happy to let the platform decide concurrency; pull when you need to cap concurrency precisely, process long jobs, or batch acknowledgements for throughput.
Step 3 — Configure a Push Subscription Correctly
For tracking, push is right; the configuration just needs an ack deadline above the handler's worst case, authentication, a dead-letter topic, and a retry policy with backoff.
gcloud pubsub subscriptions create tracking-push \
--topic=shipment-scanned \
--push-endpoint=https://tracking-abc123-ew.a.run.app/pubsub \
--push-auth-service-account=pubsub-pusher@my-project.iam.gserviceaccount.com \
--ack-deadline=30 \ # seconds Pub/Sub waits for the HTTP response
--min-retry-delay=5s --max-retry-delay=300s \
--dead-letter-topic=shipment-scanned-dlq \
--max-delivery-attempts=10
# Cloud Run handler: the body wraps the message; 2xx acks, anything else nacks
import base64, json
from fastapi import FastAPI, Request, Response
app = FastAPI()
@app.post("/pubsub")
async def receive(request: Request) -> Response:
envelope = await request.json()
msg = envelope["message"]
event = json.loads(base64.b64decode(msg["data"]))
attempt = int(envelope.get("deliveryAttempt", 1)) # present when a DLQ is configured
try:
await update_tracking(event) # idempotent upsert keyed by scan id
except TransientError:
return Response(status_code=503) # nack: retried with backoff
return Response(status_code=204) # ack
With Cloud Run, also cap --max-instances and --concurrency so a burst cannot scale the service past what its database can take. That is the push-model equivalent of the concurrency cap in processing SQS with AWS Lambda.
Step 4 — Configure Streaming Pull with Flow Control
For analytics, the worker pulls with flow control so that at most eight messages are outstanding per worker process. The client library holds leases on those messages and extends their ack deadlines up to a maximum while they are processed.
# analytics_worker.py
from concurrent.futures import ThreadPoolExecutor
from google.cloud import pubsub_v1
subscriber = pubsub_v1.SubscriberClient()
path = subscriber.subscription_path("my-project", "analytics-pull")
def callback(message: pubsub_v1.subscriber.message.Message) -> None:
try:
enrich_with_route(json.loads(message.data)) # up to ~20 s, calls mapping API
message.ack()
except RateLimited:
message.nack() # redelivered after retry-policy backoff
except Exception:
message.nack()
flow = pubsub_v1.types.FlowControl(
max_messages=8, # hard cap on in-flight per process
max_bytes=16 * 1024 * 1024,
max_lease_duration=600, # stop extending after 10 minutes
)
future = subscriber.subscribe(path, callback=callback, flow_control=flow,
scheduler=pubsub_v1.subscriber.scheduler.ThreadScheduler(
ThreadPoolExecutor(max_workers=8)))
future.result() # block; cancel() on SIGTERM
Three worker replicas give an exact ceiling of 24 concurrent mapping calls, which the team sizes to the API's limit. Scale replicas on subscription/num_undelivered_messages with a maximum replica count derived from the same limit. The general rate-limiting approach is in rate limiting third-party API calls from workers.
Step 5 — Stop Double Charges with Exactly-Once Delivery
Pub/Sub's default is at-least-once: a message can be redelivered even after a successful ack if the ack is lost or arrives after the deadline. For billing, enable exactly-once delivery on a pull subscription. With it, Pub/Sub will not redeliver a message once an ack succeeds, and the client reports whether each ack actually succeeded.
gcloud pubsub subscriptions create billing-pull \
--topic=shipment-scanned \
--enable-exactly-once-delivery \
--ack-deadline=60 \
--dead-letter-topic=shipment-scanned-dlq --max-delivery-attempts=10
from google.cloud.pubsub_v1.subscriber import exceptions as sub_exceptions
def billing_callback(message) -> None:
with db.transaction():
if already_billed(message.message_id): # still keep a cheap guard
message.ack()
return
compute_and_record_surcharge(json.loads(message.data), message.message_id)
try:
message.ack_with_response().result() # confirms the ack was accepted
except sub_exceptions.AcknowledgeError as e:
log.warning("ack failed; message may be redelivered", extra={"code": e.error_code})
Exactly-once delivery covers the transport — it prevents Pub/Sub redelivering an acknowledged message within the subscription. It does not cover a crash after compute_and_record_surcharge commits but before the ack, so keep the idempotency check keyed by message id, as described in idempotent consumers with Postgres unique constraints. Exactly-once delivery is available for pull subscriptions only and adds some latency.
Step 6 — Wire Dead-Letter Topics and Watch Them
A dead-letter topic receives messages that exceeded max-delivery-attempts (between 5 and 100). Unlike SQS, it is a topic, so you need a subscription on it to keep and inspect the messages.
PROJECT_NUMBER=$(gcloud projects describe my-project --format='value(projectNumber)')
SA="service-${PROJECT_NUMBER}@gcp-sa-pubsub.iam.gserviceaccount.com"
gcloud pubsub topics add-iam-policy-binding shipment-scanned-dlq --member="serviceAccount:$SA" --role=roles/pubsub.publisher
gcloud pubsub subscriptions add-iam-policy-binding analytics-pull --member="serviceAccount:$SA" --role=roles/pubsub.subscriber
gcloud pubsub subscriptions create shipment-scanned-dlq-hold --topic=shipment-scanned-dlq \
--message-retention-duration=7d
The IAM bindings are the step most often missed: without them, messages exceed their attempts and are not forwarded, and the subscription keeps retrying. Alert on subscription/num_undelivered_messages for the hold subscription, in line with alerting on dead-letter queue growth.
Verification
# Oldest unacked message age per subscription should stay near the handler's duration
gcloud monitoring time-series list \
--filter='metric.type="pubsub.googleapis.com/subscription/oldest_unacked_message_age"' \
--interval-start-time="$(date -u -d '-15 min' +%FT%TZ)"
# Publish a poison message and confirm it lands on the DLQ hold subscription
gcloud pubsub topics publish shipment-scanned --message='{"broken": true}'
gcloud pubsub subscriptions pull shipment-scanned-dlq-hold --limit=1 --auto-ack
For analytics, check the mapping API's request metrics: concurrent requests should never exceed replicas × max_messages.
Gotchas & Edge Cases
Ack deadline vs handler time on push. A push handler slower than the ack deadline is nacked even if it succeeds, then retried — duplicate work plus wasted API calls. This was analytics' original failure.
Push endpoints and cold starts. A push burst to a scaled-to-zero Cloud Run service returns errors during cold starts, triggering push backoff and retries. Keep a minimum instance count for latency-sensitive push subscriptions.
Ordering keys reduce parallelism. Enabling message ordering delivers messages with the same ordering key in order, one at a time; a failing message blocks its key until it succeeds or is dead-lettered.
Subscription expiration. Subscriptions with no activity expire after 31 days by default. Set --expiration-period=never for subscriptions that may be idle but must keep existing.
FAQ
Is pull more expensive than push? Pricing is by data volume, not delivery model, so direct cost is similar. Pull needs always-on workers, while push can scale to zero; that compute cost is usually the real difference.
Can one subscription use both push and pull? A subscription has one delivery type at a time, though you can change it. For different consumers, create separate subscriptions.
How does Pub/Sub compare with Cloud Tasks for background jobs? Pub/Sub fans one event out to many subscribers; Cloud Tasks targets one handler per task with explicit rate limits and scheduling. See Google Cloud Tasks for HTTP workers.
Related
- Managed Cloud Queues for Background Jobs — delivery models across providers.
- Google Cloud Tasks for HTTP Workers — push with explicit rate limits.
- Competing Consumers vs Pub/Sub Fan-Out — the pattern behind subscriptions.
- Preventing Duplicate Job Execution with Idempotency — the consumer-side half of exactly-once.