Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is data skew?

PySpark · Joins & Performance

What is data skew?

Hardpyspark-23
skewperformancejoins

Question

What is data skew in Spark, and how do you detect it?

Solution

Data skew means a few keys (or partitions) hold a disproportionate share of the data. Tasks processing those keys become stragglers while other tasks sit idle.

Where it shows up

  • Joins on hot keys (user_id = NULL, country = 'US', a viral product_id)
  • Aggregations on popular keys
  • Exploding arrays unevenly

Diagram

Partitions after hash(key):
  P0 ####
  P1 #
  P2 #
  P3 ############################   <-- straggler
  P4 ##

Detection

  • Spark UI: task duration max >> median; huge shuffle read on one task
  • Pre-check key frequencies:
from pyspark.sql import functions as F

df.groupBy("user_id").count().orderBy(F.desc("count")).show(20)

Impact

Long stages, timeouts, disk spill, executor OOM, wasted cluster spend.

Interview tip

Distinguish partition count skew (uneven file sizes) from key skew (hash collisions on hot keys). Fixes differ slightly but both create stragglers.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext