Accepted answer
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Team is split on approach A vs B. Looking for experiences from similar scale.
Context: Covering index worth it for wide dimension table?
Happy to share schema snippets or metrics if useful.
Accepted answer
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Start with the execution plan, numbers beat guesses.
Tried that, partial improvement but lag still spikes on redeploy.
Note that merge on Delta still needs unique keys defined correctly.
Document the grain decision, most BI bugs turn out to be grain bugs.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
What batch size or interval worked for you at similar scale?
In our case the root cause was an implicit cast preventing pushdown.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Agree on the staging table swap. Atomic promote prevented partial reads.
Document the grain decision, most BI bugs turn out to be grain bugs.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.