Accepted answer
We saw the same issue, fixing the partition filter dropped runtime 60%.
Hitting a wall in prod and looking for patterns others have used. Minimal repro below but happy to share more context.
Context: CDC from OLTP to lakehouse, Debezium vs DMS
Happy to share schema snippets or metrics if useful.
Accepted answer
We saw the same issue, fixing the partition filter dropped runtime 60%.
Prefer a staging table plus validation gate before promoting to prod tables.
We saw the same issue, fixing the partition filter dropped runtime 60%.
Small nit: the broadcast hint gets ignored once the table is over threshold, check the UI to confirm.
Simple way to think about it: caching is a bet that you'll ask the same question again soon. If you don't, you're just paying rent on memory for nothing.
Document the grain decision, most BI bugs turn out to be grain bugs.
Prefer a staging table plus validation gate before promoting to prod tables.
Start with the execution plan, numbers beat guesses.
Consider DuckDB or Polars for this size before spinning up a cluster.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.