Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Schema-on-read vs schema-on-write

File formats & storage · Extra High-Value

Schema-on-read vs schema-on-write

Easyformat-16
schema-on-readschema-on-writelakewarehousecontracts

Question

What is the difference between schema-on-read and schema-on-write?

Solution

Schema-on-write validates and applies a schema when data is ingested. Bad records fail early (or land in a quarantine). Warehouses and typed pipelines lean this way.

Schema-on-read stores data more loosely (JSON, text, raw files) and interprets structure when you query. Classic data lakes leaned this way for flexibility.

Schema-on-write:
  source -> validate schema -> typed Parquet/Delta table -> query

Schema-on-read:
  source -> dump raw files -> query engine guesses/applies schema at read

Trade-offs

| | Schema-on-write | Schema-on-read | |-----------|------------------------------|-------------------------------| | Ingest | Stricter, more upfront work | Fast landing, flexible | | Query | Predictable types | Surprises, cast errors later | | Quality | Fail early | Fail late (dashboards break) |

Modern practice

Lakes still land raw zones, but curated Bronze→Silver→Gold (or Medallion) layers move toward schema-on-write with contracts. Table formats enforce schemas at write when configured.

Interview tip: Do not treat schema-on-read as "always wrong." Say raw landing can be flexible; analytics tables should be enforced.

PreviousNext