Batch processes data in chunks on a schedule: hourly files, nightly warehouse loads, daily aggregates.
Streaming processes events continuously (or in micro-batches) as they arrive: Kafka consumers, Flink jobs, Spark Structured Streaming.
batch: [==== hour of events ====] -> job -> table updated streaming: e->e->e->e->e->e-> ... -> continuously updated sink
Trade-offs
- Batch: simpler ops, cheaper at huge volumes, natural for reconciliations; higher latency
- Streaming: low latency, good for alerts/CDC/realtime features; more moving parts (state, watermarks, lag)
Many platforms mix both
CDC into Kafka (stream), land to lake every few minutes, dbt models hourly (batch).
Interview tip: Pick based on *latency need* and *complexity budget*, not fashion. Ask "how fresh must this be?"