PySpark data engineering interview problem. Difficulty: beginner. Pattern: Filtering. About 10 minutes. Part of the Pro drill bank.
Turn a comma-separated string column into one row per trimmed value. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
tickets has id and tags, a comma-separated string such as 'a, b,c'. Tags can have spaces around them, be blank, or contain empty items ('x,,y'). tags can also be NULL. Return one row per ticket and tag, with the spaces around each tag removed. Rows for blank, empty and NULL tags are not returned. Columns: id, tag. Assign the DataFrame to result.
Input: tickets id | tags 1 | a, b,c 2 | 3 | NULL 4 | red ,blue 5 | x,,y Output: id | tag 1 | a 1 | b 1 | c 4 | red 4 | blue 5 | x 5 | y Ticket 1 gives a, b and c. Tickets 2 (spaces only) and 3 (NULL) give nothing. Ticket 4 gives red and blue without spaces, and ticket 5 gives x and y.
Topics: lakebench, pyspark, split, explode, trim.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Turn a comma-separated string column into one row per trimmed value.
`tickets` has `id` and `tags`, a comma-separated string such as `'a, b,c'`. Tags can have spaces around them, be blank, or contain empty items (`'x,,y'`). `tags` can also be NULL. Return one row per ticket and tag, with the spaces around each tag removed. Rows for blank, empty and NULL tags are not returned. Columns: `id`, `tag`. Assign the DataFrame to `result`.