Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Row Number for Deduplication

PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Deduplication. About 16 minutes. Part of the Pro drill bank.

Keep the earliest row per user_id + event_type with row_number. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

Using row_number over (user_id, event_type) ordered by ts, keep rn == 1. Drop the helper column. Order by user_id, event_type. Assign result.

Requirements

  • Order partitions by ts ascending.
  • First event per user_id/event_type.

Examples

Input: duplicate keyed events Output: user_id | event_type | amount | ts 1 | purchase | 40.0 | 2024-01-01 10:00:00 1 | view | 10.0 | 2024-01-01 09:00:00 2 | purchase | 50.0 | 2024-01-02 08:00:00 Earliest ts wins inside each partition.

Topics: lakebench, pyspark, row_number, window.

More PySpark interview questions · All interview problems · Learn data engineering

intermediate

Row Number for Deduplication

Interview-style drill: Keep the earliest row per user_id + event_type with row_number.

Using `row_number` over (`user_id`, `event_type`) ordered by `ts`, keep rn == 1. Drop the helper column. Order by `user_id`, `event_type`. Assign `result`.