Both change partition count, but they move data differently.
repartition(n) / repartition(cols...)
- Can increase or decrease partitions
- Full shuffle for even redistribution
- Use before large writes when you need balanced files, or to partition by key
coalesce(n)
- Typically used to reduce partitions
- Avoids a full shuffle by merging existing partitions (narrow dependency)
- Can leave uneven sizes if you shrink aggressively
# Evenly redistribute to 200 partitions (shuffle) df2 = df.repartition(200) # Partition by key (shuffle) df3 = df.repartition(100, "user_id") # Reduce to 10 partitions without full shuffle df4 = df.coalesce(10)
Diagram
repartition(4): any -> shuffle -> balanced 4 coalesce(2): P0+P1 merge, P2+P3 merge (no full reshuffle)
When to use which
- Writing many tiny outputs after a big filter:
coalesce(or AQE coalesce) carefully. - Skewed / uneven partitions needing balance:
repartition. - Never
coalesce(1)on huge data unless you truly need a single file and can afford it.
Related
repartition by column is about in-memory partitioning for the next stage; partitionBy on write is about directory layout on disk.