Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Drop Duplicates

PySpark data engineering interview problem. Difficulty: beginner. Pattern: Deduplication. About 12 minutes. Free to practice.

Keep the first row per user_id + event_type. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

From df, drop duplicates on (user_id, event_type) keeping the first occurrence. Order by user_id, event_type. Assign result.

Requirements

  • One row per user_id/event_type pair.

Constraints

  • Subset is user_id + event_type.
  • Keep first occurrence.

Examples

Input: duplicate events Output: user_id | event_type | ts 1 | purchase | 2024-01-01 10:00:00 2 | view | 2024-01-01 09:00:00 3 | purchase | 2024-01-02 08:00:00 Later duplicate timestamps for the same key are discarded.

Topics: lakebench, pyspark, dropDuplicates.

More PySpark interview questions · All interview problems · Learn data engineering

beginner

Drop Duplicates

Interview-style drill: Keep the first row per user_id + event_type.

From `df`, drop duplicates on (`user_id`, `event_type`) keeping the first occurrence. Order by `user_id`, `event_type`. Assign `result`.