Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. coalesce vs repartition

PySpark · DataFrame API & I/O

coalesce vs repartition

Mediumpyspark-18
coalescerepartitionpartitionsperformance

Question

What is the difference between coalesce and repartition in PySpark?

Solution

Both change partition count, but they move data differently.

repartition(n) / repartition(cols...)

  • Can increase or decrease partitions
  • Full shuffle for even redistribution
  • Use before large writes when you need balanced files, or to partition by key

coalesce(n)

  • Typically used to reduce partitions
  • Avoids a full shuffle by merging existing partitions (narrow dependency)
  • Can leave uneven sizes if you shrink aggressively
# Evenly redistribute to 200 partitions (shuffle)
df2 = df.repartition(200)

# Partition by key (shuffle)
df3 = df.repartition(100, "user_id")

# Reduce to 10 partitions without full shuffle
df4 = df.coalesce(10)

Diagram

repartition(4):  any -> shuffle -> balanced 4
coalesce(2):     P0+P1 merge, P2+P3 merge (no full reshuffle)

When to use which

  • Writing many tiny outputs after a big filter: coalesce (or AQE coalesce) carefully.
  • Skewed / uneven partitions needing balance: repartition.
  • Never coalesce(1) on huge data unless you truly need a single file and can afford it.

Related

repartition by column is about in-memory partitioning for the next stage; partitionBy on write is about directory layout on disk.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext