PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Window Functions. About 16 minutes. Part of the Pro drill bank.
Show rank, dense_rank, and row_number per user by amount desc. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
For each user_id, order by amount desc then ts and add rank_r, dense_r, row_r. Select those plus user_id, amount, ts. Order by user_id, row_r. Assign result.
Input: tied amounts for user 1 Output: user_id | amount | ts | rank_r | dense_r | row_r 1 | 40.0 | 2024-01-01 10:00:00 | 1 | 1 | 1 1 | 40.0 | 2024-01-01 11:00:00 | 1 | 1 | 2 1 | 10.0 | 2024-01-01 09:00:00 | 3 | 2 | 3 Each user's eligible order totals are aggregated; users with higher totals appear earlier in the ranked output.
Topics: lakebench, pyspark, rank, dense_rank, row_number.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Show rank, dense_rank, and row_number per user by amount desc.
For each `user_id`, order by `amount` desc then `ts` and add `rank_r`, `dense_r`, `row_r`. Select those plus `user_id`, `amount`, `ts`. Order by `user_id`, `row_r`. Assign `result`.