Accepted answer
Use a staging table plus swap pattern: load to stg, validate, then atomic partition replace. That avoids any partially-visible state.
Loading a daily sales snapshot into the warehouse. Team is debating full partition overwrite versus merge on the natural key.
Failures mid-run leave partial data with overwrite. Merge is slower but safer. What actually drives the decision for you?
Accepted answer
Use a staging table plus swap pattern: load to stg, validate, then atomic partition replace. That avoids any partially-visible state.
Consider DuckDB or Polars for this size before spinning up a cluster.
Note that merge on Delta still needs unique keys defined correctly.
In our case the root cause was an implicit cast preventing pushdown.
Start with the execution plan, numbers beat guesses.
Overwrite the partition when the data is a complete replacement for that day and reruns are cheap. Merge when updates for the same keys can still arrive within the window.
Our dimension is slowly changing, does that change the join order?
In our case the root cause was an implicit cast preventing pushdown.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.