Pandas data engineering interview problem. Difficulty: beginner. Pattern: Filtering. About 10 minutes. Part of the Pro drill bank.
Turn a comma-separated string column into one trimmed row per value. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
df has id and tags, a comma-separated string such as 'a, b ,c'. Values can have spaces around them, be empty ('x,,y'), or be missing entirely. Return one row per id and tag with the spaces removed. Empty and missing tags are not returned. Columns: id, tag. Keep the original order of ids and of tags inside each id, and reset the index. Assign the DataFrame to result.
Input: df id | tags 1 | a, b ,c 2 | 3 | NULL 4 | red ,blue 5 | x,,y Output: id | tag 1 | a 1 | b 1 | c 4 | red 4 | blue 5 | x 5 | y Ticket 1 gives a, b and c. Row 2 (spaces only) and row 3 (missing) give nothing; 4 gives red and blue; 5 gives x and y.
Topics: lakebench, pandas, str.split, explode, strip.
More interview problems · All interview problems · Learn data engineering
Interview-style drill: Turn a comma-separated string column into one trimmed row per value.
`df` has `id` and `tags`, a comma-separated string such as `'a, b ,c'`. Values can have spaces around them, be empty (`'x,,y'`), or be missing entirely. Return one row per `id` and tag with the spaces removed. Empty and missing tags are not returned. Columns: `id`, `tag`. Keep the original order of ids and of tags inside each id, and reset the index. Assign the DataFrame to `result`.