In pull-based ingestion, your pipeline goes and asks the source for data on a schedule. In push-based ingestion, the source sends data to you when something happens. The difference is who controls the timing.
Pull
A job calls an API or runs a query against a database every hour, and takes what is new.
scheduler -> pipeline -> source system (API / database)
You control the pace, the retries and the load you put on the source, and if your system is down, nothing is lost: you catch up on the next run, because the data is still at the source. The downside is latency. Data is only as fresh as your schedule. Polling a lot to get lower latency wastes calls, may hit rate limits, and puts load on the source.
Push
The source sends events, webhooks or files to you as they happen.
source -> HTTP endpoint / queue / bucket -> pipeline
Latency is low, and nothing is fetched when nothing changes. But now the source sets the pace. A burst (10,000 webhooks in a minute) hits you at once, and if your endpoint is down, the sender may retry for a while and then give up, so events are lost. You also must deal with duplicates (senders retry), out-of-order events, and authentication of the sender.
Make push reliable: put a queue in front
Accept the message with a very small, fast endpoint that writes it into a durable queue or log (Kafka, SQS, Pub/Sub, or object storage), and returns success. Process from the queue at your own speed. This absorbs bursts, survives downstream outages, and lets you replay.
Choosing
- The source supports events and you need low latency: push, with a queue.
- You need control, the source only offers an API, or hourly freshness is enough: pull.
- Many systems combine them: webhooks for fast updates and a periodic pull to catch anything missed. That belt and braces design is common for SaaS sources.
Add that with pull you should use incremental watermarks, and with push you must deduplicate by an event id.