Apache Parquet is a columnar, binary file format designed for analytics. It stores typed columns, compresses them, and embeds schema plus per-chunk statistics so engines can skip useless data.
Problem first. CSV and JSON are easy for humans but painful at scale: weak typing, poor compression, and you usually decode whole rows even when you need two columns.
How Parquet is laid out (mental model)
Parquet file ├── footer (schema, offsets) └── row groups └── column chunks └── pages (compressed data + optional stats)
Why lakes prefer it
- Column pruning: read
amountwithout touchingnotes - Predicate pushdown: skip row groups when min/max cannot match
amount > 100 - Compression: Snappy, ZSTD, etc. shrink storage and scan time
- Splittable: Spark/Hive can parallelize reads of large files
- Schema travels with the file
Interview tip: Lead with "columnar + compression + stats." Mention it is the default analytical format under Delta/Iceberg/Hudi data files.