Accepted answer
Start with the execution plan, numbers beat guesses.
Team is split on approach A vs B. Looking for experiences from similar scale.
Context: How many files per partition is too many for Parquet?
Happy to share schema snippets or metrics if useful.
Accepted answer
Start with the execution plan, numbers beat guesses.
Yes, the partial index was the win for us too.
How do you handle backfill without duplicating rows?
Start with the execution plan, numbers beat guesses.
Small nit: the broadcast hint gets ignored once the table is over threshold, check the UI to confirm.
This matches our runbook for skewed keys.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Note that merge on Delta still needs unique keys defined correctly.
In our case the root cause was an implicit cast preventing pushdown.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
In our case the root cause was an implicit cast preventing pushdown.
Could you share a sketch of the salting logic?
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.