Accepted answer
We saw the same issue, fixing the partition filter dropped runtime 60%.
I have read the docs but real-world tradeoffs are unclear. What would you optimize first?
Context: CDC from OLTP to lakehouse, Debezium vs DMS
Happy to share schema snippets or metrics if useful.
Accepted answer
We saw the same issue, fixing the partition filter dropped runtime 60%.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
We replaced custom sensors with data contracts and row count checks.
Yes, the partial index was the win for us too.
Start with the execution plan, numbers beat guesses.
Worth measuring the serialized size before choosing broadcast.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
In our case the root cause was an implicit cast preventing pushdown.
Note that merge on Delta still needs unique keys defined correctly.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.