A watermark is a threshold that tells Spark how late events are still allowed for a given event-time column. It bounds state for event-time windows so the engine can drop too-late data and shrink aggregation state.
Why needed
With event-time windows, data can arrive late or out of order. Without a watermark, Spark must keep state forever for every window key.
Event time progress -----------------------> watermark = max_event_time_seen - allowed_lateness Events older than watermark may be dropped for stateful aggs
Example
from pyspark.sql import functions as F
windowed = (
events
.withWatermark("event_time", "10 minutes")
.groupBy(F.window("event_time", "5 minutes"), "user_id")
.count()
)Meaning: once Spark believes event time has moved on, events more than 10 minutes late may be ignored for this stateful aggregation.
Interview tip
Distinguish processing time vs event time, and say watermarks are about state bounding and late data policy, not about sources magically becoming ordered.