Python data engineering interview problem. Difficulty: beginner. Pattern: Strings. About 8 minutes. Part of the Pro drill bank.
Build pairs of consecutive words after lowercasing and removing punctuation. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
Implement bigrams(text: str) -> list[tuple]. Lowercase the text, remove every punctuation character (such as , . ! ? ' -), split it into words on whitespace and return the pairs of consecutive words as tuples, in order. A token that is only punctuation disappears. Fewer than two words gives an empty list.
Input: bigrams("Hello, World! Hello") Output: [('hello', 'world'), ('world', 'hello')] After lowercasing and removing punctuation the words are hello, world, hello.
Topics: lakebench, python, tokenize, n-grams.
More Python interview questions · All interview problems · Learn data engineering
Interview-style drill: Build pairs of consecutive words after lowercasing and removing punctuation.
Implement `bigrams(text: str) -> list[tuple]`. Lowercase the text, remove every punctuation character (such as `, . ! ? ' -`), split it into words on whitespace and return the pairs of consecutive words as tuples, in order. A token that is only punctuation disappears. Fewer than two words gives an empty list. Example: `"The cat sat."` gives `[('the', 'cat'), ('cat', 'sat')]`.