Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Handle a source schema change

Pipelines & scenarios · Design & Scenarios

Handle a source schema change

Hardpipe-22
schema evolutioncontractsdriftbackfillcompatibility

Question

How do you handle a source schema change in a data pipeline?

Solution

Sources add, rename, or type-change columns. Pipelines must detect drift, avoid silent corruption, and evolve curated models safely.

Producer adds orders.discount_code
        |
Ingestion lands new field in bronze (good: keep raw flexible)
        |
Contract / schema check
  - alert if unexpected break (type change, dropped column)
  - allow additive fields if policy says so
        |
Silver/gold models updated
  - dbt model + tests
  - backfill if historical values needed

Practical playbook

1. Land raw as-is (JSON or wide bronze) so you do not lose new fields. 2. Detect drift: schema registry, Great Expectations, dbt tests, contract CI. 3. Classify the change: additive (usually safe) vs breaking (rename/drop/type). 4. Version contracts with producers when possible. 5. Evolve silver/gold explicitly; do not auto-remap renames blindly. 6. Backfill if BI needs history for the new column.

Breaking change example

amount string → number. Gate the deploy, fix casts, validate samples, then promote.

Interview tip: Separate "absorb in raw" from "consciously evolve curated." Mention contracts and alerts so schema changes are not discovered by broken dashboards.

PreviousNext