PySpark data engineering interview problem. Difficulty: advanced. Pattern: Explain Plans. About 20 minutes. Part of the Pro drill bank.
Shrink partitions with coalesce(2) and return the ordered frame. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Call coalesce(2) on df (prints a note in the shim), then select all columns ordered by id. Prefer coalesce over repartition when only reducing partition count. Assign result.
Input: partitioned-ish fixture Output: id | date | region | amount 1 | 2024-01-01 | US | 10.0 2 | 2024-01-01 | EU | 20.0 3 | 2024-01-02 | US | 30.0 4 | 2024-01-02 | EU | 40.0 5 | 2024-01-03 | US | 50.0 6 | 2024-01-03 | APAC | 60.0 Data unchanged; coalesce only affects partition count.
Topics: lakebench, pyspark, coalesce, repartition.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Shrink partitions with coalesce(2) and return the ordered frame.
Call `coalesce(2)` on `df` (prints a note in the shim), then select all columns ordered by `id`. Prefer coalesce over repartition when only reducing partition count. Assign `result`.