Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Broken job: driver OOM from .collect()

PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Driver Memory. About 20 minutes. Part of the Pro drill bank.

Production ticket: Spark driver died after collecting the clickstream to build silver events. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

The PySpark silver job 'worked' on a 5k-row sample. Friday's 40M-row replay OOMd the driver because the intern called .collect(). Fix the script: spark.table / parquet read, no .collect() or .toPandas() one row per event_id, latest ingest_time same columns as the SQL silver contract, ordered by event_id The incident log has the executor dump. Treat this as a production RCA, not a LeetCode puzzle.

Requirements

  • A given event_id may appear more than once (replay writes); output keeps exactly one row per event_id at the latest ingest_time.
  • Output columns and ordering match the SQL silver contract: event_id, user_id, session_id, event_type, product_id, device, country, event_time, ordered by event_id.
  • Silver clickstream without a driver collect; unique event_id.

Examples

Input: event_id | user_id | ingest_time E00013 | U053 | 2026-08-07 09:10:08 E00013 | U053 | 2026-08-07 09:29:08 Output: -- rows.collect() + spark.createDataFrame(rows) is a pure pass-through: it does -- no dedup at all, so both replay rows for E00013 would still be there. -- In production this line never finishes: collecting 40M rows OOMs the -- driver (container killed, code 137) before any output is produced. The bug is not just the OOM: even ignoring memory, .collect() followed by createDataFrame() performs no deduplication, so it can't satisfy the one-row-per-event_id contract either.

Topics: pyspark, OOM, collect.

More PySpark interview questions · All interview problems · Learn data engineering

intermediate

Broken job: driver OOM from .collect()

Production ticket: Spark driver died after collecting the clickstream to build silver events.

The PySpark silver job 'worked' on a 5k-row sample. Friday's 40M-row replay OOMd the driver because the intern called `.collect()`. Fix the script: - spark.table / parquet read, **no .collect() or .toPandas()** - one row per event_id, latest ingest_time - same columns as the SQL silver contract, ordered by event_id The incident log has the executor dump. Treat this as a production RCA, not a LeetCode puzzle.