PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Schema Drift. About 16 minutes. Part of the Pro drill bank.
Add a missing phone column as null. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
df should expose user_id, email, phone but phone is missing. Add phone as null. Order by user_id. Assign result.
Input: users without phone Output: user_id | email | phone 1 | a@x.com | None 2 | b@x.com | None 3 | c@x.com | None Missing column added as null.
Topics: lakebench, pyspark, withColumn, lit.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Add a missing phone column as null.
`df` should expose `user_id`, `email`, `phone` but `phone` is missing. Add `phone` as null. Order by `user_id`. Assign `result`.