Accepted answer
Start with the execution plan, numbers beat guesses.
Every tuning guide says "don't leave it at 200" but picking the actual number feels like folklore. How do you arrive at a real value instead of guessing?
Accepted answer
Start with the execution plan, numbers beat guesses.
Yes, the partial index was the win for us too.
In our case the root cause was an implicit cast preventing pushdown.
If you're on Spark 3.x with AQE enabled, spark.sql.adaptive.coalescePartitions.enabled lets Spark pick a reasonable number after the fact, which is more robust than one fixed static setting across jobs of varying size.
Note that merge on Delta still needs unique keys defined correctly.
Rule of thumb: target 100 to 200MB per shuffle partition. Look at the shuffle write size for your stage in the Spark UI and divide by that target to get a partition count.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.