Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is Parquet?

File formats & storage · Formats & Table Formats

What is Parquet?

Easyformat-02
Parquetcolumnaranalyticslake

Question

What is Apache Parquet, and why do data lakes use it?

Solution

Apache Parquet is a columnar, binary file format designed for analytics. It stores typed columns, compresses them, and embeds schema plus per-chunk statistics so engines can skip useless data.

Problem first. CSV and JSON are easy for humans but painful at scale: weak typing, poor compression, and you usually decode whole rows even when you need two columns.

How Parquet is laid out (mental model)

Parquet file                                                ├── footer (schema, offsets)                              └── row groups                                                  └── column chunks                                               └── pages (compressed data + optional stats)

Why lakes prefer it

  • Column pruning: read amount without touching notes
  • Predicate pushdown: skip row groups when min/max cannot match amount > 100
  • Compression: Snappy, ZSTD, etc. shrink storage and scan time
  • Splittable: Spark/Hive can parallelize reads of large files
  • Schema travels with the file

Interview tip: Lead with "columnar + compression + stats." Mention it is the default analytical format under Delta/Iceberg/Hudi data files.

PreviousNext