PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Aggregation. About 14 minutes. Part of the Pro drill bank.
Per user, the distinct items as a sorted array, their count, and whether a gift card is among them. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
purchases has user and item; item can be NULL and users buy the same item more than once. Return one row per user with: items: the distinct items as an array sorted alphabetically (NULL items are ignored) item_count: the number of distinct items has_gift_card: true if the array contains 'gift_card' Columns: user, items, item_count, has_gift_card. Assign the DataFrame to result.
Input: purchases user | item u1 | book u1 | pen u1 | book u1 | gift_card u2 | pen u2 | NULL u3 | mug u3 | mug Output: user | items | item_count | has_gift_card u1 | ["book","gift_card","pen"] | 3 | True u2 | ["pen"] | 1 | False u3 | ["mug"] | 1 | False u1 bought book twice but it is listed once; u2's NULL item is ignored; u3's repeated mug gives a single item.
Topics: lakebench, pyspark, collect_set, array, array_contains.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Per user, the distinct items as a sorted array, their count, and whether a gift card is among them.
`purchases` has `user` and `item`; `item` can be NULL and users buy the same item more than once. Return one row per user with: - `items`: the distinct items as an array sorted alphabetically (NULL items are ignored) - `item_count`: the number of distinct items - `has_gift_card`: true if the array contains `'gift_card'` Columns: `user`, `items`, `item_count`, `has_gift_card`. Assign the DataFrame to `result`.