Apache Avro is a row-oriented binary format with a rich schema system. It is popular for event streams, Kafka payloads, and write-heavy ingest paths where you append whole records.
Avro strengths
- Compact binary encoding of full records
- Excellent schema evolution story (writer/reader schemas, compatibility rules)
- Fast to serialize/deserialize whole events
- Common with schema registries (Confluent, Glue, etc.)
Parquet strengths
- Columnar scans for analytics
- Column pruning and row-group skipping
- Better compression for warehouse-style queries
Typical path:
producers -> Avro events (Kafka / landing zone)
-> convert/compact to Parquet (lake / warehouse tables)Avro vs Parquet (interview table)
| Dimension | Avro | Parquet | |------------------|------------------------------|------------------------------| | Layout | Row-oriented | Columnar | | Best for | Streaming, write/append | Analytic reads | | Schema evolution | Very strong | Supported, more limited | | Scan 2 of 50 cols| Still pays more row cost | Reads only those columns |
Interview tip: "Avro for the pipe; Parquet for the lake." Do not claim one replaces the other; they solve different stages.