Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Repartition vs coalesce

Batch & Streaming · Batch Processing

Repartition vs coalesce

Mediumstream-19
repartitioncoalescepartitionsshuffle

Question

What is the difference between repartition and coalesce in Spark?

Solution

Both change the number of partitions, but they move data differently.

`repartition(n)` does a full shuffle to roughly even partitions. Use to increase parallelism or rebalance after a filter. Optional column args: repartition(n, "user_id").

`coalesce(n)` reduces partitions by merging existing ones without a full shuffle (when shrinking). Faster, but partitions may be uneven. Great before writing fewer output files.

repartition(8):  any -> shuffle -> 8 balanced partitions
coalesce(2):     8 partitions -> merge locally -> ~2 (no full exchange)

Rules of thumb

  • Scale up parallelism → repartition
  • Scale down before write → coalesce
  • Do not coalesce(1) on huge data unless you accept a single-writer bottleneck

Interview tip: "repartition shuffles; coalesce shrinks cheaply." Mention output file count as a common reason to coalesce.

PreviousNext