PySpark data engineering interview problem. Difficulty: beginner. Pattern: Aggregation. About 12 minutes. Part of the Pro drill bank.
Count how often each word appears in lines of text, ignoring case and punctuation. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
lines is a DataFrame with one column, line, holding raw text. Some lines are empty and some contain punctuation. Return the frequency of every word: lowercase the text, remove punctuation, split on whitespace and ignore empty tokens. Columns: word, count. Most frequent first, then alphabetical order for equal counts. Assign the DataFrame to result.
Input: lines line Spark makes big data simple Big data, big wins! SPARK spark Output: word | count big | 3 spark | 3 data | 2 makes | 1 simple | 1 wins | 1 spark appears three times once case is ignored, big three times (with the punctuation removed), data twice, and the empty line adds nothing.
Topics: lakebench, pyspark, word count, split, groupby.
More PySpark interview questions · All interview problems · Learn data engineering
Interview-style drill: Count how often each word appears in lines of text, ignoring case and punctuation.
`lines` is a DataFrame with one column, `line`, holding raw text. Some lines are empty and some contain punctuation. Return the frequency of every word: lowercase the text, remove punctuation, split on whitespace and ignore empty tokens. Columns: `word`, `count`. Most frequent first, then alphabetical order for equal counts. Assign the DataFrame to `result`.