PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Filtering. About 16 minutes. Part of the Pro drill bank.
Build full_name with concat_ws instead of a multi-arg UDF. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Add full_name from first_name and last_name with concat_ws. Keep all three columns. Order by first_name. Assign result.
Input: names Output: first_name | last_name | full_name ada | lovelace | ada lovelace alan | turing | alan turing grace | hopper | grace hopper Native concat avoids UDF serialization.
Topics: lakebench, pyspark, concat_ws.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Build full_name with concat_ws instead of a multi-arg UDF.
Add `full_name` from `first_name` and `last_name` with `concat_ws`. Keep all three columns. Order by `first_name`. Assign `result`.