In our case the root cause was an implicit cast preventing pushdown.
Team is split on approach A vs B. Looking for experiences from similar scale.
Context: Remote DE roles, async communication tips for ICs
Happy to share schema snippets or metrics if useful.
In our case the root cause was an implicit cast preventing pushdown.
Another path: push the compute to the warehouse if the data's already there.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Agree on the staging table swap. Atomic promote prevented partial reads.
Worth measuring the serialized size before choosing broadcast.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Worth measuring the serialized size before choosing broadcast.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.