PySpark data engineering interview problem. Difficulty: beginner. Pattern: Deduplication. About 12 minutes. Free to practice.
Keep the first row per user_id + event_type. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
From df, drop duplicates on (user_id, event_type) keeping the first occurrence. Order by user_id, event_type. Assign result.
Input: duplicate events Output: user_id | event_type | ts 1 | purchase | 2024-01-01 10:00:00 2 | view | 2024-01-01 09:00:00 3 | purchase | 2024-01-02 08:00:00 Later duplicate timestamps for the same key are discarded.
Topics: lakebench, pyspark, dropDuplicates.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Keep the first row per user_id + event_type.
From `df`, drop duplicates on (`user_id`, `event_type`) keeping the first occurrence. Order by `user_id`, `event_type`. Assign `result`.