PySpark data engineering interview problem. Difficulty: advanced. Pattern: Medallion. About 20 minutes. Part of the Pro drill bank.
coalesce(1) before a simulated write to avoid tiny files. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Too many output files usually means too many partitions. Call coalesce(1) on df, then select all columns ordered by id (stand-in for a single-file write). Assign result.
Input: six rows Output: id | date | region | amount 1 | 2024-01-01 | US | 10.0 2 | 2024-01-01 | EU | 20.0 3 | 2024-01-02 | US | 30.0 4 | 2024-01-02 | EU | 40.0 5 | 2024-01-03 | US | 50.0 6 | 2024-01-03 | APAC | 60.0 Rows unchanged after coalesce(1).
Topics: lakebench, pyspark, coalesce.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: coalesce(1) before a simulated write to avoid tiny files.
Too many output files usually means too many partitions. Call `coalesce(1)` on `df`, then select all columns ordered by `id` (stand-in for a single-file write). Assign `result`.