Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Repartition vs Coalesce

PySpark data engineering interview problem. Difficulty: advanced. Pattern: Explain Plans. About 20 minutes. Part of the Pro drill bank.

Shrink partitions with coalesce(2) and return the ordered frame. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

Call coalesce(2) on df (prints a note in the shim), then select all columns ordered by id. Prefer coalesce over repartition when only reducing partition count. Assign result.

Constraints

  • Call coalesce(2).
  • Do not drop rows.

Examples

Input: partitioned-ish fixture Output: id | date | region | amount 1 | 2024-01-01 | US | 10.0 2 | 2024-01-01 | EU | 20.0 3 | 2024-01-02 | US | 30.0 4 | 2024-01-02 | EU | 40.0 5 | 2024-01-03 | US | 50.0 6 | 2024-01-03 | APAC | 60.0 Data unchanged; coalesce only affects partition count.

Topics: lakebench, pyspark, coalesce, repartition.

More PySpark interview questions · All interview problems · Learn data engineering

advanced

Repartition vs Coalesce

Interview-style drill: Shrink partitions with coalesce(2) and return the ordered frame.

Call `coalesce(2)` on `df` (prints a note in the shim), then select all columns ordered by `id`. Prefer coalesce over repartition when only reducing partition count. Assign `result`.