PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Window Functions. About 16 minutes. Part of the Pro drill bank.
Top 3 amounts per user_id via row_number. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
For each user_id, keep the top 3 rows by amount descending using row_number. Return user_id, amount, rn. Order by user_id, rn. Assign result.
Input: amounts per user Output: user_id | amount | rn 1 | 40.0 | 1 1 | 30.0 | 2 1 | 20.0 | 3 2 | 50.0 | 1 2 | 5.0 | 2 User 1 drops the 10.0 fourth place; user 2 has only two rows.
Topics: lakebench, pyspark, row_number, top-n.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Top 3 amounts per user_id via row_number.
For each `user_id`, keep the top 3 rows by `amount` descending using `row_number`. Return `user_id`, `amount`, `rn`. Order by `user_id`, `rn`. Assign `result`.