Schema evolution is how a dataset's schema changes over time while old and new data remain usable: add columns, drop fields, widen types, rename carefully, or mark columns optional.
v1: {order_id, amount}
v2: {order_id, amount, currency} # additive, usually safe
v3: {order_id, amount_cents} # rename/type change, riskyHow different layers handle it
- Avro + schema registry: explicit compatibility modes (backward, forward, full)
- Parquet datasets: readers may fill missing columns with nulls; type changes are tricky
- Delta / Iceberg / Hudi: table metadata tracks schema versions; engines enforce or evolve with rules
Safer practices
- Prefer additive, nullable columns
- Avoid silent renames; version contracts
- Dual-write / dual-read for breaking changes, then backfill
- Test consumers against compatibility policy in CI
Interview tip: Classify changes as additive vs breaking, then mention table-format metadata and (for streams) registry compatibility. Schema evolution is a process, not a checkbox.