Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. File formats: Parquet vs CSV

Data platform · Platform Concepts

File formats: Parquet vs CSV

Mediumplatform-23
ParquetCSVcolumnarlake

Question

Why do data lakes prefer Parquet (or columnar formats) over CSV?

Solution

CSV is simple text: easy to eyeball, terrible for large analytics. Parquet (and ORC) stores data in columns with compression and embedded schema/types.

CSV row store mindset: read whole lines even if you need 2 columns
Parquet columns:      read only amount + region files/chunks

Why Parquet wins at scale

  • Columnar reads → less I/O for analytic queries
  • Compression → cheaper storage and faster scans
  • Typed schema → fewer silent string bugs
  • Works well with Spark, Hive, BigQuery external tables, etc.

CSV still fine for

Tiny exchanges, human debugging, some partner feeds (then convert on ingest).

Interview tip: Lead with columnar projection + compression. Mention schema and predicate pushdown as extras.

PreviousNext