Late data arrives after you already processed the time window it belongs to. Example: an event with event_time=10:01 shows up at 12:40 because a mobile app was offline.
window 10:00-11:00 closed at 11:05 late event for 10:15 arrives at 12:40 -> needs a policy
Common strategies
- Watermarks (streaming): wait a grace period, then close the window
- Allowed lateness: update aggregates for a while after the window
- Recompute partitions: nightly job rebuilds yesterday with all arrivals
- Separate late path: write corrections / adjustments instead of rewriting everything
Batch angle
Even daily jobs see late data: a file for ds=2026-09-04 arriving on Sept 6. Idempotent partition rebuilds or merge-on-key fix this.
Interview tip: Distinguish event time vs processing time, then give one streaming (watermark) and one batch (partition rebuild) answer.