Streaming processing means handling events continuously (or in tiny micro-batches) as they arrive, instead of waiting for a daily file.
event → event → event → event → ...
\ | | /
continuous job (Flink / Spark Structured Streaming / Kafka consumer)
|
near-real-time sink (DB, topic, dashboard)Typical use cases
- Fraud / anomaly alerts within seconds
- CDC into a lake or search index
- Real-time feature updates for ML
- Live ops dashboards (lag of seconds/minutes)
What gets harder
State, late data, watermarks, exactly-once semantics, consumer lag, and ops complexity.
Hybrid reality
Many platforms stream into Kafka or a lake, then run batch dbt models on a schedule. Streaming and batch are tools, not religions.
Interview tip: Pick streaming when latency is a product requirement. Otherwise prefer batch for simplicity. Say "how fresh must this be?" before naming Kafka.