Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Compression vs encoding

File formats & storage · Performance & Layout

Compression vs encoding

Mediumformat-14
compressionencodingdictionaryParquetORC

Question

What is the difference between compression and encoding in columnar formats?

Solution

In Parquet/ORC, encoding and compression are related but not the same step.

Encoding chooses a compact *representation* of column values based on data patterns, still typed and often CPU-cheap to decode:

  • Dictionary encoding for low-cardinality strings (status: PAID/SHIPPED/…)
  • Run-length encoding (RLE) for repeated values
  • Delta / bit-packing for integers and timestamps

Compression then applies a general codec (Snappy, ZSTD, gzip, LZ4) to encoded byte pages to shrink size further.

raw column values
    -> encoding (dictionary / RLE / delta)
    -> compression codec (Snappy / ZSTD / ...)
    -> stored page bytes

Why both matter

  • Encoding exploits column structure (repetition, limited domains)
  • Compression exploits remaining byte redundancy
  • Snappy: fast, moderate ratio; ZSTD: often better ratio; gzip: smaller but slower (and historically worse for splits in some stacks)

Interview tip: "Encoding is format-smart representation; compression is a general codec on the encoded bytes." Columnar layout makes both much more effective than compressing whole CSV rows.

PreviousNext