In Parquet/ORC, encoding and compression are related but not the same step.
Encoding chooses a compact *representation* of column values based on data patterns, still typed and often CPU-cheap to decode:
- Dictionary encoding for low-cardinality strings (
status: PAID/SHIPPED/…) - Run-length encoding (RLE) for repeated values
- Delta / bit-packing for integers and timestamps
Compression then applies a general codec (Snappy, ZSTD, gzip, LZ4) to encoded byte pages to shrink size further.
raw column values
-> encoding (dictionary / RLE / delta)
-> compression codec (Snappy / ZSTD / ...)
-> stored page bytesWhy both matter
- Encoding exploits column structure (repetition, limited domains)
- Compression exploits remaining byte redundancy
- Snappy: fast, moderate ratio; ZSTD: often better ratio; gzip: smaller but slower (and historically worse for splits in some stacks)
Interview tip: "Encoding is format-smart representation; compression is a general codec on the encoded bytes." Columnar layout makes both much more effective than compressing whole CSV rows.