PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Filtering. About 16 minutes. Part of the Pro drill bank.
Keep only user_id, email, phone. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
From a wide df, keep user_id, email, phone only. Order by user_id. Assign result.
Input: wide users Output: user_id | email | phone 1 | a@x.com | 111 2 | b@x.com | 222 3 | c@x.com | 333 Extra columns removed.
Topics: lakebench, pyspark, select.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Keep only user_id, email, phone.
From a wide `df`, keep `user_id`, `email`, `phone` only. Order by `user_id`. Assign `result`.