ZSTD is the modern standard recommendation for Parquet because it matches or exceeds GZIP compression ratios while decompressing at speeds close to Snappy. Snappy focuses on low CPU overhead and fast decompression speed with a moderate compression ratio, making it the historic default for active compute pipelines. GZIP achieves higher space savings than Snappy but burns substantial CPU cycles during both compression and decompression, making it less attractive for interactive analytics.
Codec comparison on CPU and ratio
Codec Ratio Compression Speed Decompression Speed Default In Snappy Moderate Very Fast Very Fast Legacy Spark / Hive ZSTD High Fast (Level 1-3) Very Fast Iceberg / DuckDB / Trino GZIP High Slow Slow Cold text archives
The trade-offs between these codecs show why engine defaults have shifted:
- Snappy was designed at Google to prioritize throughput over disk savings, ensuring that CPU cores spend time executing queries rather than decompressing blocks.
- ZSTD, developed by Meta, provides configurable compression levels (commonly level 1 to level 3 in data engines), offering exceptional space reduction without crippling decompression throughput.
- GZIP is CPU-heavy and generally obsolete for modern columnar Parquet storage, though it still surfaces in legacy Hadoop architectures.
Why text compression hurts split execution
When dealing with uncompressed raw text files like CSV or JSON, compression codec choice introduces an operational trap:
- GZIP compresses files as a single continuous LZ77 stream with dynamic Huffman trees. Because decompression state depends on preceding bytes, a query engine cannot split a single 20 GB .csv.gz file across multiple executors; one CPU core must decompress the entire file from start to finish.
- BZIP2 provides block markers every 900 KB that permit distributed splitting, but its decompression algorithms are exceedingly slow on CPU, making it impractical for high-throughput pipelines.
Production codec selection
For modern analytical tables, configure ZSTD as your default Parquet compression codec. It reduces cloud storage bills, lowers network transfer bytes between object stores and compute clusters, and decompresses quickly enough that CPU rarely becomes the bottleneck.