Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Group-by on CSV text without pandas

Python data engineering interview problem. Difficulty: intermediate. Pattern: Dictionaries. About 16 minutes. Part of the Pro drill bank.

Average an amount per region from CSV text using only the csv module, counting bad rows instead of failing. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

Implement avg_by_region(csv_text: str) -> tuple. csv_text has a header line region,amount followed by data rows. Use only the standard library csv module, not pandas. Return (averages, bad_rows): averages: dict of region to the mean amount, rounded to 2 decimals, regions sorted by name bad_rows: how many rows could not be used A row is bad when it does not have exactly two fields, its region is blank, or its amount is not a finite number. Bad rows are counted, never raised. Completely blank lines are skipped and are not counted. Regions are matched after trimming spaces.

Requirements

  • Rounded to 2 decimal places.
  • Blank lines are not bad rows.

Constraints

  • Standard library only.
  • Up to 100,000 rows.
  • The header is always present.

Examples

Input: avg_by_region("region,amount\n\"Asia, East\",10\nEU,5\nEU,15\n") Output: ({'Asia, East': 10.0, 'EU': 10.0}, 0) The quoted region keeps its comma. EU averages 5 and 15 to 10.0.

Topics: lakebench, python, csv, group by, bad rows.

More Python interview questions · All interview problems · Learn data engineering

intermediate

Group-by on CSV text without pandas

Interview-style drill: Average an amount per region from CSV text using only the csv module, counting bad rows instead of failing.

Implement `avg_by_region(csv_text: str) -> tuple`. `csv_text` has a header line `region,amount` followed by data rows. Use only the standard library `csv` module, not pandas. Return `(averages, bad_rows)`: - `averages`: dict of region to the mean `amount`, rounded to 2 decimals, regions sorted by name - `bad_rows`: how many rows could not be used A row is **bad** when it does not have exactly two fields, its region is blank, or its amount is not a finite number. Bad rows are counted, never raised. Completely blank lines are skipped and are not counted. Regions are matched after trimming spaces.