Accepted answer
In our case the root cause was an implicit cast preventing pushdown.
Team is split on approach A vs B. Looking for experiences from similar scale.
Context: Covering index worth it for wide dimension table?
Happy to share schema snippets or metrics if useful.
Accepted answer
In our case the root cause was an implicit cast preventing pushdown.
What batch size or interval worked for you at similar scale?
Small nit: the broadcast hint gets ignored once the table is over threshold, check the UI to confirm.
Worth measuring the serialized size before choosing broadcast.
Prefer a staging table plus validation gate before promoting to prod tables.
Start with the execution plan, numbers beat guesses.
Quick plain-English version: the database is doing more work than it needs to because it can't tell in advance which rows actually match. An index is basically a shortcut list so it doesn't have to check every single row.
+1, saw identical behaviour after upgrading Spark 3.4 to 3.5.
We saw the same issue, fixing the partition filter dropped runtime 60%.
Small nit: the broadcast hint gets ignored once the table is over threshold, check the UI to confirm.
We saw the same issue, fixing the partition filter dropped runtime 60%.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.