PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Schema Drift. About 16 minutes. Part of the Pro drill bank.
Report columns added/removed between two snapshots. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Compare old and new column sets. Return one row with added_columns (comma-joined new-only names), removed_columns (comma-joined old-only names), and type_changes set to 'none' for this fixture. Assign result.
Input: old vs new Output: added_columns | removed_columns | type_changes size | | none size appears only on new.
Topics: lakebench, pyspark, schema.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Report columns added/removed between two snapshots.
Compare `old` and `new` column sets. Return one row with `added_columns` (comma-joined new-only names), `removed_columns` (comma-joined old-only names), and `type_changes` set to `'none'` for this fixture. Assign `result`.