PySpark data engineering interview problem. Difficulty: advanced. Pattern: Streaming. About 20 minutes. Part of the Pro drill bank.
Count events in 5-minute buckets (no stream window() API). Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Bucket event_time into 5-minute starts 10:00, 10:05, 10:10 and count events per window_start as event_count. Order by window_start. Assign result.
Input: event times Output: window_start | event_count 2024-01-01 10:00:00 | 2 2024-01-01 10:05:00 | 2 2024-01-01 10:10:00 | 2 Two events land in each 5-minute bucket.
Topics: lakebench, pyspark, window, groupBy.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Count events in 5-minute buckets (no stream window() API).
Bucket `event_time` into 5-minute starts `10:00`, `10:05`, `10:10` and count events per `window_start` as `event_count`. Order by `window_start`. Assign `result`.