Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Batch vs streaming

Data platform · Pipelines

Batch vs streaming

Easyplatform-02
batchstreaminglatencyarchitecture

Question

What is the difference between batch and streaming data pipelines?

Solution

Batch processes data in chunks on a schedule: hourly files, nightly warehouse loads, daily aggregates.

Streaming processes events continuously (or in micro-batches) as they arrive: Kafka consumers, Flink jobs, Spark Structured Streaming.

batch:     [==== hour of events ====] -> job -> table updated
streaming: e->e->e->e->e->e-> ...     -> continuously updated sink

Trade-offs

  • Batch: simpler ops, cheaper at huge volumes, natural for reconciliations; higher latency
  • Streaming: low latency, good for alerts/CDC/realtime features; more moving parts (state, watermarks, lag)

Many platforms mix both

CDC into Kafka (stream), land to lake every few minutes, dbt models hourly (batch).

Interview tip: Pick based on *latency need* and *complexity budget*, not fashion. Ask "how fresh must this be?"

PreviousNext