Schema-on-write validates and applies a schema when data is ingested. Bad records fail early (or land in a quarantine). Warehouses and typed pipelines lean this way.
Schema-on-read stores data more loosely (JSON, text, raw files) and interprets structure when you query. Classic data lakes leaned this way for flexibility.
Schema-on-write: source -> validate schema -> typed Parquet/Delta table -> query Schema-on-read: source -> dump raw files -> query engine guesses/applies schema at read
Trade-offs
| | Schema-on-write | Schema-on-read | |-----------|------------------------------|-------------------------------| | Ingest | Stricter, more upfront work | Fast landing, flexible | | Query | Predictable types | Surprises, cast errors later | | Quality | Fail early | Fail late (dashboards break) |
Modern practice
Lakes still land raw zones, but curated Bronze→Silver→Gold (or Medallion) layers move toward schema-on-write with contracts. Table formats enforce schemas at write when configured.
Interview tip: Do not treat schema-on-read as "always wrong." Say raw landing can be flexible; analytics tables should be enforced.