Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Word count

PySpark data engineering interview problem. Difficulty: beginner. Pattern: Aggregation. About 12 minutes. Part of the Pro drill bank.

Count how often each word appears in lines of text, ignoring case and punctuation. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.

lines is a DataFrame with one column, line, holding raw text. Some lines are empty and some contain punctuation. Return the frequency of every word: lowercase the text, remove punctuation, split on whitespace and ignore empty tokens. Columns: word, count. Most frequent first, then alphabetical order for equal counts. Assign the DataFrame to result.

Requirements

  • Output columns are word and count.
  • Sort by count descending, then word.

Constraints

  • Words are separated by whitespace.
  • Punctuation is anything that is not a letter, digit or whitespace.

Examples

Input: lines line Spark makes big data simple Big data, big wins! SPARK spark Output: word | count big | 3 spark | 3 data | 2 makes | 1 simple | 1 wins | 1 spark appears three times once case is ignored, big three times (with the punctuation removed), data twice, and the empty line adds nothing.

Topics: lakebench, pyspark, word count, split, groupby.

More PySpark interview questions · All interview problems · Learn data engineering

beginner

Word count

Interview-style drill: Count how often each word appears in lines of text, ignoring case and punctuation.

`lines` is a DataFrame with one column, `line`, holding raw text. Some lines are empty and some contain punctuation. Return the frequency of every word: lowercase the text, remove punctuation, split on whitespace and ignore empty tokens. Columns: `word`, `count`. Most frequent first, then alphabetical order for equal counts. Assign the DataFrame to `result`.