Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Deduplicate Stream

PySpark data engineering interview problem. Difficulty: advanced. Pattern: Deduplication. About 20 minutes. Part of the Pro drill bank.

dropDuplicates on event_id for a stream-like frame. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

Deduplicate events on event_id keeping the first row. Order by event_id. Assign result.

Requirements

  • Unique event_id rows.

Constraints

  • Subset event_id.
  • Keep first occurrence.

Examples

Input: duplicate event ids Output: event_id | user_id | event_time E1 | 1 | 2024-01-01 10:00:00 E2 | 2 | 2024-01-01 10:02:00 E3 | 3 | 2024-01-01 10:03:00 First occurrence per event_id wins.

Topics: lakebench, pyspark, dropDuplicates, streaming.

More PySpark interview questions · All interview problems · Learn data engineering

advanced

Deduplicate Stream

Interview-style drill: dropDuplicates on event_id for a stream-like frame.

Deduplicate `events` on `event_id` keeping the first row. Order by `event_id`. Assign `result`.