Batch processes a finite pile of data after it has landed: last night's files, yesterday's table, an hour of events. You start a job, it reads the whole chunk, writes results, and stops.
Streaming processes an unbounded flow of events as they arrive (or nearly so). The job keeps running. Latency is seconds or milliseconds, not hours.
Problem first. An e-commerce company needs nightly revenue reports *and* fraud alerts within seconds of a weird payment. One schedule cannot serve both. Batch fits the report. Streaming fits the alert.
BATCH (finite) [day of events] ---> job runs once ---> table updated ---> done STREAM (unbounded) e -> e -> e -> e -> ... ---> always-on job ---> sink keeps updating
End-to-end DE mental model (how batch and stream sit in a real platform):
Sources Ingest Process Serve
-------- ------ ------- -----
OLTP DBs --CDC--> Kafka / files
APIs --poll--> object storage ---> batch Spark/dbt ---> warehouse
Apps --push--> landing zone stream Flink/SS dashboards
Logs --------> (windows, joins) APIs / ML
^-- orchestration (Airflow) schedules batch
^-- stream jobs run continuously with checkpointsTrade-offs interviewers want
- Batch: simpler ops, great for huge reprocesses and reconciliations, higher latency
- Stream: low latency, more state/watermarks/lag to manage
- Most platforms mix both: stream ingest + batch transforms
Interview tip: Do not say "streaming is always better." Ask how fresh the data must be, then pick batch, micro-batch, or true streaming.