Sometimes the real fix is a product change so you stop needing that join at all.
Currently using Snappy for all Parquet output by default. Considering switching to zstd for a better compression ratio, but worried about CPU cost on both write and read paths.
Sometimes the real fix is a product change so you stop needing that join at all.
Note that merge on Delta still needs unique keys defined correctly.
Careful with NULL in join keys, they'll drop rows in an inner join.
Benchmarked on our largest table, zstd cut storage about 35% with negligible read latency impact. Rolling it out as the new default.
Prefer a staging table plus validation gate before promoting to prod tables.
zstd typically gives noticeably better compression than snappy for similar CPU cost on modern hardware, especially at moderate compression levels like level 3. It's a reasonable default switch for most analytical workloads.
Start with the execution plan, numbers beat guesses.
What batch size or interval worked for you at similar scale?
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.