Sometimes the real fix is a product change so you stop needing that join at all.
Files land in a folder, a script picks them up, transforms, and loads. If the script crashes mid-run and retries, some rows can get double-loaded. Is chasing true exactly-once worth it here, or is idempotent at-least-once enough?
Sometimes the real fix is a product change so you stop needing that join at all.
Consider DuckDB or Polars for this size before spinning up a cluster.
We saw the same issue, fixing the partition filter dropped runtime 60%.
Start with the execution plan, numbers beat guesses.
Worth measuring the serialized size before choosing broadcast.
Event-driven beats cron once landing time gets unpredictable.
Note that merge on Delta still needs unique keys defined correctly.
At-least-once plus an idempotent load, merge or upsert on a natural key, or delete-then-insert per source file, gets you the same practical outcome as exactly-once without the complexity of distributed transactions.
This matches our runbook for skewed keys.
This matches our runbook for skewed keys.
In our case the root cause was an implicit cast preventing pushdown.
Could you share a sketch of the salting logic?
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.