Row-oriented storage keeps all columns of one row together on disk. Columnar storage keeps each column together across many rows.
Problem first. Analytics queries often touch a few columns over millions of rows (SELECT region, SUM(amount)). Transactional systems often need the whole row for one key (SELECT * FROM orders WHERE id = 42). The layout on disk should match the access pattern.
Row store (one row = contiguous bytes): [id=1, name=A, amount=10] [id=2, name=B, amount=20] ... Column store (one column = contiguous bytes): id: [1, 2, ...] name: [A, B, ...] amount: [10, 20, ...]
Why columnar wins for lakes / warehouses
- Read only the columns you need (less I/O)
- Similar values compress better when packed together
- Engines can vectorize scans and apply stats (min/max) per column chunk
Why row wins for OLTP
- One seek/read fetches the whole record
- Updates to a single row are localized
Interview tip: "Row for whole-record lookups; column for analytic scans of a few fields across many rows." Name Parquet/ORC as columnar and Avro/JSON/CSV as typically row-ish.