Pandas data engineering interview problem. Difficulty: beginner. Pattern: Aggregation. About 12 minutes. Free to practice.
Aggregate chunked frames without loading one giant CSV. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Treat chunks as successive read_csv(..., chunksize=) pieces. Sum amount per user_id across chunks into result with columns user_id, amount, sorted by user_id. Do not concatenate raw chunks before aggregating if you can aggregate per chunk first.
Input: three chunks Output: user_id | amount 1 | 15.0 2 | 21.0 3 | 7.0 4 | 9.0 Chunk-local sums merge correctly.
Topics: lakebench, pandas, chunks, groupby.
More interview problems · All interview problems · Learn data engineering
Interview-style drill: Aggregate chunked frames without loading one giant CSV.
Treat `chunks` as successive `read_csv(..., chunksize=)` pieces. Sum `amount` per `user_id` across chunks into `result` with columns `user_id`, `amount`, sorted by `user_id`. Do not concatenate raw chunks before aggregating if you can aggregate per chunk first.