PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Window Functions. About 16 minutes. Part of the Pro drill bank.
percent_rank of amount within each user_id. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
For each user_id, compute percentile_rank of amount ascending. Return user_id, amount, percentile_rank ordered by user_id, amount. Assign result.
Input: three amounts per user Output: user_id | amount | percentile_rank 1 | 10.0 | 0.0 1 | 20.0 | 0.5 1 | 30.0 | 1.0 percent_rank is (rank-1)/(n-1) inside the partition.
Topics: lakebench, pyspark, percent_rank.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: percent_rank of amount within each user_id.
For each `user_id`, compute `percentile_rank` of `amount` ascending. Return `user_id`, `amount`, `percentile_rank` ordered by `user_id`, `amount`. Assign `result`.