Event time is when the event *actually happened* in the real world (phone click, payment authorized, sensor reading).
Processing time is when your pipeline *saw* the event (Kafka consumer read it, Spark task handled it).
They diverge when networks delay, clocks skew, or jobs pause.
User clicks at 10:00:00 --> event_time = 10:00:00
Event arrives at 10:00:07 --> processing_time = 10:00:07
(7s late / out-of-order risk)Tiny example. Hourly unique visitors by *click time*:
- Processing-time window: a late click from 9:59 counted in the 10:00 bucket → wrong report
- Event-time window: still attributed to 9:00–10:00 after it arrives late
timeline
event_time: |-- window 9-10 --|-- window 10-11 --|
late event @9:58 arrives at processing_time 10:05
should still land in window 9-10Interview tip: "Business correctness uses event time. Ops dashboards often use processing time." Watermarks exist because event time is messy.