Batch processing means collecting data over a time window, then processing it as one chunk on a schedule (hourly, daily, weekly).
00:00 -------- events accumulate -------- 01:00
|
batch job
|
table updatedCommon shapes
- Nightly warehouse load of yesterday's orders
- Hourly Spark job compacting Kafka dumps into Parquet
- Weekly backfill of a corrected metric
Why teams still love batch
- Simpler ops than streaming (clear start/end, easy to reason about)
- Cheap at huge volumes (process once per window)
- Natural for reconciliations and financial close
Cost of batch
Latency. Dashboards may be hours stale. That is fine for many BI use cases and wrong for fraud alerts.
Interview tip: "Batch = process a closed window of data on a schedule." Contrast with streaming only after you define the latency need.