Python data engineering interview problem. Difficulty: intermediate. Pattern: Strings. About 14 minutes. Part of the Pro drill bank.
Parse log lines, ignore malformed and non-error lines, and rank error codes with a stable tie-break. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Implement top_errors(lines: list[str], n: int) -> list[tuple]. A log line looks like "2024-05-01 10:00:00 ERROR E500 database timeout": date, time, level, an error code, then a free-text message (the message may be missing). Count error codes over lines whose level is exactly ERROR and whose code matches E followed by three digits. Lines at other levels, and lines that are malformed (too few fields, or a bad code), are skipped without raising an error. Return the n most frequent codes as (code, count) tuples, most frequent first. Codes with equal counts are ordered by code, ascending. If there are fewer than n codes, return all of them.
Input: top_errors(["2024-05-01 10:00:00 ERROR E500 db timeout", "2024-05-01 10:01:00 ERROR E404 not found", "2024-05-01 10:02:00 ERROR E500 db timeout", "2024-05-01 10:03:00 ERROR E404 again", "2024-05-01 10:04:00 INFO I100 started"], 2) Output: [('E404', 2), ('E500', 2)] E404 and E500 both appear twice, so code order decides: E404 first. The INFO line is ignored.
Topics: lakebench, python, logs, parsing, ranking.
More Python interview questions · All interview problems · Learn data engineering
Interview-style drill: Parse log lines, ignore malformed and non-error lines, and rank error codes with a stable tie-break.
Implement `top_errors(lines: list[str], n: int) -> list[tuple]`. A log line looks like `"2024-05-01 10:00:00 ERROR E500 database timeout"`: date, time, level, an error code, then a free-text message (the message may be missing). Count error codes over lines whose level is exactly `ERROR` and whose code matches `E` followed by three digits. Lines at other levels, and lines that are malformed (too few fields, or a bad code), are skipped without raising an error. Return the `n` most frequent codes as `(code, count)` tuples, most frequent first. Codes with equal counts are ordered by code, ascending. If there are fewer than `n` codes, return all of them.