Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is data skew?

Pipelines & scenarios · Core Pipeline Concepts

What is data skew?

Mediumpipe-10
skewstragglerssaltingjoinsperformance

Question

What is data skew, and how do you handle it in pipelines?

Solution

Data skew means a few keys hold a disproportionate share of the data. Tasks on those keys become stragglers while other tasks finish and sit idle.

join / groupBy key = country

  US:     ############################  (huge)
  IN:     ##########
  DE:     ##
  IS:     #

One executor stuck on US; cluster looks "busy" but progress crawls.

Where skew shows up

  • Joins and aggregations on popular keys (null, unknown, big customers)
  • Hot partitions in Kafka
  • Uneven file sizes after bad partitioning

Mitigations

1. Filter / isolate hot keys and process them separately 2. Salting: add a random suffix to hot keys, then aggregate twice 3. Broadcast join the small side instead of shuffling both 4. Adaptive skew join (Spark AQE) when available 5. Repartition with a better key or more partitions

Interview tip: Define skew as uneven key distribution causing stragglers, then name salting or broadcast as fixes. Mention checking Spark stage timelines for one long task.

PreviousNext