PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Aggregation. About 16 minutes. Part of the Pro drill bank.
Each user's spend as a percent of all spend. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Sum amount per user_id as total_spend, then compute pct_of_total = round(100 * total_spend / grand_total, 2). Order by user_id. Assign result.
Input: spend rows Output: user_id | total_spend | pct_of_total 1 | 40.0 | 36.36 2 | 20.0 | 18.18 3 | 50.0 | 45.45 Grand total is 110.
Topics: lakebench, pyspark, ratio, sum.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Each user's spend as a percent of all spend.
Sum `amount` per `user_id` as `total_spend`, then compute `pct_of_total` = round(100 * total_spend / grand_total, 2). Order by `user_id`. Assign `result`.