CSV is simple text: easy to eyeball, terrible for large analytics. Parquet (and ORC) stores data in columns with compression and embedded schema/types.
CSV row store mindset: read whole lines even if you need 2 columns Parquet columns: read only amount + region files/chunks
Why Parquet wins at scale
- Columnar reads → less I/O for analytic queries
- Compression → cheaper storage and faster scans
- Typed schema → fewer silent string bugs
- Works well with Spark, Hive, BigQuery external tables, etc.
CSV still fine for
Tiny exchanges, human debugging, some partner feeds (then convert on ingest).
Interview tip: Lead with columnar projection + compression. Mention schema and predicate pushdown as extras.