Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Share of group and above-group-mean

Pandas data engineering interview problem. Difficulty: intermediate. Pattern: GroupBy. About 14 minutes. Part of the Pro drill bank.

Add each order's share of its customer's total and a flag for being above the customer's mean, without collapsing rows. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

df has one row per order: order_id, customer, amount. Keep every row and add: share: the order's amount divided by that customer's total amount, rounded to 4 decimals above_mean: True when the order's amount is strictly greater than the mean amount of that customer's orders Columns: order_id, customer, amount, share, above_mean. Sort by order_id, reset the index. Assign the DataFrame to result.

Requirements

  • Same number of rows as the input.

Constraints

  • amount is positive.
  • Every customer has at least one order.

Examples

Input: df order_id | customer | amount 1 | a | 10 2 | a | 30 3 | a | 20 4 | b | 50 5 | b | 50 6 | c | 80 Output: order_id | customer | amount | share | above_mean 1 | a | 10 | 0.1667 | False 2 | a | 30 | 0.5 | True 3 | a | 20 | 0.3333 | False 4 | b | 50 | 0.5 | False 5 | b | 50 | 0.5 | False 6 | c | 80 | 1 | False Customer a's total is 60, so order 2 is half (0.5) and above the mean of 20. Customer b's two equal orders are not above their mean. Customer c has one order: share 1 and not above its own mean.

Topics: lakebench, pandas, transform, groupby, share.

More interview problems · All interview problems · Learn data engineering

intermediate

Share of group and above-group-mean

Interview-style drill: Add each order's share of its customer's total and a flag for being above the customer's mean, without collapsing rows.

`df` has one row per order: `order_id`, `customer`, `amount`. Keep every row and add: - `share`: the order's `amount` divided by that customer's total amount, rounded to 4 decimals - `above_mean`: `True` when the order's `amount` is strictly greater than the mean amount of that customer's orders Columns: `order_id`, `customer`, `amount`, `share`, `above_mean`. Sort by `order_id`, reset the index. Assign the DataFrame to `result`.