Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Choosing a compression codec

File formats & storage · Parquet Internals

Choosing a compression codec

Easyfile-formats-23
compressionsnappyzstdgzipparquet

Question

Snappy, GZIP or ZSTD: which compression would you choose for Parquet and why?

Solution

ZSTD is the modern standard recommendation for Parquet because it matches or exceeds GZIP compression ratios while decompressing at speeds close to Snappy. Snappy focuses on low CPU overhead and fast decompression speed with a moderate compression ratio, making it the historic default for active compute pipelines. GZIP achieves higher space savings than Snappy but burns substantial CPU cycles during both compression and decompression, making it less attractive for interactive analytics.

Codec comparison on CPU and ratio

Codec    Ratio     Compression Speed   Decompression Speed   Default In
Snappy   Moderate  Very Fast           Very Fast             Legacy Spark / Hive
ZSTD     High      Fast (Level 1-3)    Very Fast             Iceberg / DuckDB / Trino
GZIP     High      Slow                Slow                  Cold text archives

The trade-offs between these codecs show why engine defaults have shifted:

  • Snappy was designed at Google to prioritize throughput over disk savings, ensuring that CPU cores spend time executing queries rather than decompressing blocks.
  • ZSTD, developed by Meta, provides configurable compression levels (commonly level 1 to level 3 in data engines), offering exceptional space reduction without crippling decompression throughput.
  • GZIP is CPU-heavy and generally obsolete for modern columnar Parquet storage, though it still surfaces in legacy Hadoop architectures.

Why text compression hurts split execution

When dealing with uncompressed raw text files like CSV or JSON, compression codec choice introduces an operational trap:

  • GZIP compresses files as a single continuous LZ77 stream with dynamic Huffman trees. Because decompression state depends on preceding bytes, a query engine cannot split a single 20 GB .csv.gz file across multiple executors; one CPU core must decompress the entire file from start to finish.
  • BZIP2 provides block markers every 900 KB that permit distributed splitting, but its decompression algorithms are exceedingly slow on CPU, making it impractical for high-throughput pipelines.

Production codec selection

For modern analytical tables, configure ZSTD as your default Parquet compression codec. It reduces cloud storage bills, lowers network transfer bytes between object stores and compute clusters, and decompresses quickly enough that CPU rarely becomes the bottleneck.

PreviousNext