Both change the number of partitions, but they move data differently.
`repartition(n)` does a full shuffle to roughly even partitions. Use to increase parallelism or rebalance after a filter. Optional column args: repartition(n, "user_id").
`coalesce(n)` reduces partitions by merging existing ones without a full shuffle (when shrinking). Faster, but partitions may be uneven. Great before writing fewer output files.
repartition(8): any -> shuffle -> 8 balanced partitions coalesce(2): 8 partitions -> merge locally -> ~2 (no full exchange)
Rules of thumb
- Scale up parallelism →
repartition - Scale down before write →
coalesce - Do not
coalesce(1)on huge data unless you accept a single-writer bottleneck
Interview tip: "repartition shuffles; coalesce shrinks cheaply." Mention output file count as a common reason to coalesce.