Schema drift is unexpected change in a dataset's structure over time: new columns, missing columns, renamed fields, type changes, or altered nullability. Pipelines that assumed yesterday's schema then fail or silently mis-map data.
Day 1 CSV: order_id, amount, currency Day 2 CSV: order_id, amount_cents, currency <-- rename/type drift
Dangerous vs safer changes
- Safer: additive optional columns
- Dangerous: renames, type changes, removed required fields, semantic changes with the same name
Detection
- Compare inferred schema vs contract / schema registry
- Alert on column set diffs in observability tools
- Fail ingestion when required fields disappear
- dbt/GX tests that assert columns exist and types match
Handling strategies
- Prefer additive evolution with versioned contracts
- Kafka schema registry compatibility modes (backward/forward/full)
- Quarantine bad batches; do not silently coerce everything
- Dual-write / dual-read migrations for breaking changes
Interview tip: Define drift as unplanned schema change, give a rename example, then mention contract checks and additive-first evolution.