PySpark data engineering interview problem. Difficulty: advanced. Pattern: Deduplication. About 20 minutes. Part of the Pro drill bank.
dropDuplicates on event_id for a stream-like frame. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Deduplicate events on event_id keeping the first row. Order by event_id. Assign result.
Input: duplicate event ids Output: event_id | user_id | event_time E1 | 1 | 2024-01-01 10:00:00 E2 | 2 | 2024-01-01 10:02:00 E3 | 3 | 2024-01-01 10:03:00 First occurrence per event_id wins.
Topics: lakebench, pyspark, dropDuplicates, streaming.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: dropDuplicates on event_id for a stream-like frame.
Deduplicate `events` on `event_id` keeping the first row. Order by `event_id`. Assign `result`.