Accepted answer
In our case the root cause was an implicit cast preventing pushdown.
I have read the docs but real-world tradeoffs are unclear. What would you optimize first?
Context: Data vault hype, when is it actually justified?
Happy to share schema snippets or metrics if useful.
Accepted answer
In our case the root cause was an implicit cast preventing pushdown.
Event-driven beats cron once landing time gets unpredictable.
Small nit: the broadcast hint gets ignored once the table is over threshold, check the UI to confirm.
Yes, the partial index was the win for us too.
Small nit: the broadcast hint gets ignored once the table is over threshold, check the UI to confirm.
Our dimension is slowly changing, does that change the join order?
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Worth measuring the serialized size before choosing broadcast.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Idempotent writes with merge keys saved us during backfills.
Start with the execution plan, numbers beat guesses.
In our case the root cause was an implicit cast preventing pushdown.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.