Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Items per user as an array

PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Aggregation. About 14 minutes. Part of the Pro drill bank.

Per user, the distinct items as a sorted array, their count, and whether a gift card is among them. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

purchases has user and item; item can be NULL and users buy the same item more than once. Return one row per user with: items: the distinct items as an array sorted alphabetically (NULL items are ignored) item_count: the number of distinct items has_gift_card: true if the array contains 'gift_card' Columns: user, items, item_count, has_gift_card. Assign the DataFrame to result.

Requirements

  • Distinct items only.
  • Sorted array.

Constraints

  • Items are strings.
  • Output array order is alphabetical.

Examples

Input: purchases user | item u1 | book u1 | pen u1 | book u1 | gift_card u2 | pen u2 | NULL u3 | mug u3 | mug Output: user | items | item_count | has_gift_card u1 | ["book","gift_card","pen"] | 3 | True u2 | ["pen"] | 1 | False u3 | ["mug"] | 1 | False u1 bought book twice but it is listed once; u2's NULL item is ignored; u3's repeated mug gives a single item.

Topics: lakebench, pyspark, collect_set, array, array_contains.

More PySpark interview questions · All interview problems · Learn data engineering

intermediate

Items per user as an array

Interview-style drill: Per user, the distinct items as a sorted array, their count, and whether a gift card is among them.

`purchases` has `user` and `item`; `item` can be NULL and users buy the same item more than once. Return one row per user with: - `items`: the distinct items as an array sorted alphabetically (NULL items are ignored) - `item_count`: the number of distinct items - `has_gift_card`: true if the array contains `'gift_card'` Columns: `user`, `items`, `item_count`, `has_gift_card`. Assign the DataFrame to `result`.