Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Handle Schema Drift

PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Schema Drift. About 16 minutes. Part of the Pro drill bank.

Report columns added/removed between two snapshots. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

Compare old and new column sets. Return one row with added_columns (comma-joined new-only names), removed_columns (comma-joined old-only names), and type_changes set to 'none' for this fixture. Assign result.

Constraints

  • One output row.
  • Join names with commas, no spaces required.

Examples

Input: old vs new Output: added_columns | removed_columns | type_changes size | | none size appears only on new.

Topics: lakebench, pyspark, schema.

More PySpark interview questions · All interview problems · Learn data engineering

intermediate

Handle Schema Drift

Interview-style drill: Report columns added/removed between two snapshots.

Compare `old` and `new` column sets. Return one row with `added_columns` (comma-joined new-only names), `removed_columns` (comma-joined old-only names), and `type_changes` set to `'none'` for this fixture. Assign `result`.