PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Driver Memory. About 20 minutes. Part of the Pro drill bank.
Production ticket: Spark driver died after collecting the clickstream to build silver events. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
The PySpark silver job 'worked' on a 5k-row sample. Friday's 40M-row replay OOMd the driver because the intern called .collect(). Fix the script: spark.table / parquet read, no .collect() or .toPandas() one row per event_id, latest ingest_time same columns as the SQL silver contract, ordered by event_id The incident log has the executor dump. Treat this as a production RCA, not a LeetCode puzzle.
Input: event_id | user_id | ingest_time E00013 | U053 | 2026-08-07 09:10:08 E00013 | U053 | 2026-08-07 09:29:08 Output: -- rows.collect() + spark.createDataFrame(rows) is a pure pass-through: it does -- no dedup at all, so both replay rows for E00013 would still be there. -- In production this line never finishes: collecting 40M rows OOMs the -- driver (container killed, code 137) before any output is produced. The bug is not just the OOM: even ignoring memory, .collect() followed by createDataFrame() performs no deduplication, so it can't satisfy the one-row-per-event_id contract either.
Topics: pyspark, OOM, collect.
More PySpark interview questions · All interview problems · Learn data engineering
Production ticket: Spark driver died after collecting the clickstream to build silver events.
The PySpark silver job 'worked' on a 5k-row sample. Friday's 40M-row replay OOMd the driver because the intern called `.collect()`. Fix the script: - spark.table / parquet read, **no .collect() or .toPandas()** - one row per event_id, latest ingest_time - same columns as the SQL silver contract, ordered by event_id The incident log has the executor dump. Treat this as a production RCA, not a LeetCode puzzle.