Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Read Large CSV in Chunks

Pandas data engineering interview problem. Difficulty: beginner. Pattern: Aggregation. About 12 minutes. Free to practice.

Aggregate chunked frames without loading one giant CSV. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

Treat chunks as successive read_csv(..., chunksize=) pieces. Sum amount per user_id across chunks into result with columns user_id, amount, sorted by user_id. Do not concatenate raw chunks before aggregating if you can aggregate per chunk first.

Constraints

  • Final columns user_id, amount.
  • Sort by user_id.

Examples

Input: three chunks Output: user_id | amount 1 | 15.0 2 | 21.0 3 | 7.0 4 | 9.0 Chunk-local sums merge correctly.

Topics: lakebench, pandas, chunks, groupby.

More interview problems · All interview problems · Learn data engineering

beginner

Read Large CSV in Chunks

Interview-style drill: Aggregate chunked frames without loading one giant CSV.

Treat `chunks` as successive `read_csv(..., chunksize=)` pieces. Sum `amount` per `user_id` across chunks into `result` with columns `user_id`, `amount`, sorted by `user_id`. Do not concatenate raw chunks before aggregating if you can aggregate per chunk first.