Accepted answer
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Hitting a wall in prod and looking for patterns others have used. Minimal repro below but happy to share more context.
Context: Incremental model full refresh strategy in prod
Happy to share schema snippets or metrics if useful.
Accepted answer
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Small nit: the broadcast hint gets ignored once the table is over threshold, check the UI to confirm.
For anyone who wants the two-sentence version: data arrived out of order, and the code assumed it wouldn't. Fix the assumption or sort the data, pick one.
Tried that, partial improvement but lag still spikes on redeploy.
We saw the same issue, fixing the partition filter dropped runtime 60%.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.