PySpark data engineering interview problem. Difficulty: intermediate. Pattern: Window Functions. About 16 minutes. Part of the Pro drill bank.
Running sum of amount per user ordered by ts. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
For each user_id, compute running_total = sum(amount) ordered by ts. Return user_id, ts, amount, running_total ordered by user_id, ts. Assign result.
Input: events ordered by ts Output: user_id | ts | amount | running_total 1 | 2024-01-01 09:00:00 | 10.0 | 10.0 1 | 2024-01-01 10:00:00 | 40.0 | 50.0 1 | 2024-01-01 11:00:00 | 40.0 | 90.0 Default sum frame accumulates through the current row.
Topics: lakebench, pyspark, sum, window.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Running sum of amount per user ordered by ts.
For each `user_id`, compute `running_total` = sum(amount) ordered by `ts`. Return `user_id`, `ts`, `amount`, `running_total` ordered by `user_id`, `ts`. Assign `result`.