Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Source adds a new column without telling you

Pipelines & scenarios · Production Scenarios

Source adds a new column without telling you

Mediumpipelines-45
scenarioschema-driftdata-contractsschema-evolutionquarantine

Question

The source system added three columns and changed one type overnight. What should your pipeline do?

Solution

Treat the two kinds of change differently. New columns are usually safe to accept. A changed type is dangerous, because it can silently corrupt data or break downstream logic. The pipeline should detect both, handle each by a clear rule, and tell a person.

Detect

Compare the incoming schema (from the file header, the API response, or the table metadata) with the expected schema stored in your repo or registry, at ingestion time, before any transformation. Do this on every run. It takes seconds.

Rules by change type

  • New columns: let them through into the raw (bronze) layer automatically, so no data is lost, and alert the owner. Do not promote them to curated tables until someone reviews what they mean and whether they hold personal data. Many formats and table engines support this evolution (for example Delta's mergeSchema, or a rescued data column in Auto Loader).
  • Type change (for example amount from integer to string): do not load blindly. Either fail the run for that table, or route the affected records to quarantine. If you cast it silently, you may get NULLs or wrong values, and nobody notices until a report looks strange.
  • Dropped or renamed columns: fail or alert. Downstream models that reference the old name will break, so it is better to stop early with a clear message.

What the pipeline should do on a failure

Keep the raw data, because losing it is the real risk. Stop only the downstream steps that depend on the changed table, and publish the previous good version of curated tables. Alert with the table name, the old schema, the new schema and the diff.

Prevent it

  • A data contract with the producer: an agreed schema, types, meaning of fields, and a rule that breaking changes are announced with notice and versioned.
  • For streams, a schema registry with compatibility rules (backward compatible only), so incompatible changes are rejected at the producer.
  • Versioned tables or views for breaking changes, so consumers migrate on their own timeline.
  • Tests in CI on the producer side, if you can get them.

Say it in the interview

Additive changes: accept in bronze, review before promoting. Breaking changes: stop or quarantine. And talk to the source team, because the long-term fix is a contract, not a smarter pipeline.

PreviousNext