Python data engineering interview problem. Difficulty: intermediate. Pattern: Hashing. About 12 minutes. Part of the Pro drill bank.
Find the comment that appears in the most locations, counting a repeat inside one location once. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Implement most_common(comments_by_location: dict) -> str | None. The input maps a location name to the list of comments posted there. A comment counts once per location, however many times it appears in that location's list. Return the comment that appears in the largest number of locations. When several comments tie, return the one that is smallest in normal string order. Return None if there are no comments at all.
Input: most_common({"a": ["good", "good", "bad"], "b": ["good", "ok"], "c": ["bad"]}) Output: 'bad' good and bad each appear in two locations (the repeat in a counts once); bad is smaller, so it wins the tie.
Topics: lakebench, python, counting, tie-break, sets.
More Python interview questions · All interview problems · Learn data engineering
Interview-style drill: Find the comment that appears in the most locations, counting a repeat inside one location once.
Implement `most_common(comments_by_location: dict) -> str | None`. The input maps a location name to the list of comments posted there. A comment counts **once per location**, however many times it appears in that location's list. Return the comment that appears in the largest number of locations. When several comments tie, return the one that is smallest in normal string order. Return `None` if there are no comments at all.