Accepted answer
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Team is split on approach A vs B. Looking for experiences from similar scale.
Context: Predicate pushdown not happening in Spark SQL join
Happy to share schema snippets or metrics if useful.
Accepted answer
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
ELI5 version: it's not broken, it's just slow because it's checking way more stuff than it needs to. Narrowing what it checks is almost always the fix.
Prefer a staging table plus validation gate before promoting to prod tables.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
In our case the root cause was an implicit cast preventing pushdown.
What batch size or interval worked for you at similar scale?
Start with the execution plan, numbers beat guesses.
Worth measuring the serialized size before choosing broadcast.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.