Micro-batch chops the stream into tiny batches (e.g. every 500ms–few seconds), runs a mini batch job per slice, then repeats. Spark Structured Streaming's classic trigger model works this way.
True / continuous streaming processes each record (or small buffer) through an always-on dataflow graph with fine-grained operators. Apache Flink is the usual example.
Micro-batch (Spark SS typical)
time ---> | batch1 | batch2 | batch3 | ...
each batch: read -> plan -> execute -> write
True streaming (Flink typical)
time ---> e->e->e->e->e->e->e-> ...
continuous operators + network buffersPractical differences
- Latency: Flink can go lower more naturally; Spark SS often second-scale (can do continuous processing in limited cases)
- API / ecosystem: Spark shines when you already use Spark batch SQL/DataFrames
- State & event time: both support windows/watermarks; Flink's streaming semantics are often considered more "native"
Interview tip: "Spark SS = micro-batches over DataFrames. Flink = continuous operators." Pick based on latency needs and team stack, not brand loyalty.