Latency and operational cost follow an inverse, non-linear relationship: true continuous streaming provides sub-second event delivery but requires dedicated, always-on compute clusters and incurs high architectural complexity. As you relax latency to micro-batches running every minute or every fifteen minutes, fixed scheduling overhead and file fragmentation decrease while hardware efficiency improves dramatically. Senior data engineers always recommend choosing the largest latency interval that business stakeholders find acceptable to maintain low operating expenses.
The three latency tiers compared
Understanding how operational dynamics shift across intervals is essential for pipeline design:
- True streaming (sub-second): Engines like Apache Flink process records individually or in tiny network buffers. Cluster nodes must remain powered on 24/7 with permanent network sockets. This delivers the lowest latency for fraud evaluation and real-time alerts, but fixed cluster hosting costs remain high even during low-traffic periods.
- One-minute micro-batch: Engines like Spark Structured Streaming trigger execution cycles every sixty seconds. At this frequency, fixed scheduling overhead (task launch, catalog resolution, state checkpointing) consumes a significant fraction of total execution time. Writing every minute also creates 1,440 tiny files per partition daily, requiring aggressive compaction.
- Fifteen-minute micro-batch: Scheduling overhead shrinks to a negligible percentage of the run. Computing tasks can run on elastic or serverless clusters that scale down during quiet hours. File writes produce naturally compact 128 MB Parquet blocks, and transaction log commits remain manageable.
The overhead penalty of tiny batches
Whenever a micro-batch interval shrinks below two minutes, the system spends more time preparing to process data than actually transforming records:
15-min batch: [-------------------- Process Data (95%) --------------------][Commit 5%] 1-min batch: [--- Schedule (40%) ---][-- Process (30%) --][--- Commit Log (30%) ---]
Every micro-batch must write transaction log entries (such as Delta Lake or Iceberg JSON commits) and register metadata. Running this cycle every sixty seconds inflates metadata catalogs and degrades downstream read performance.
Practical selection strategy
Always challenge requests for sub-minute latency by evaluating the consumers of the data. If the destination is a Tableau dashboard or a daily reporting model, running one-minute batches wastes cloud budget. Choose the fifteen-minute or hourly interval unless automated downstream actions strictly demand immediate processing.