Data skew means a few keys (or partitions) hold a disproportionate share of the data. Tasks processing those keys become stragglers while other tasks sit idle.
Where it shows up
- Joins on hot keys (
user_id = NULL,country = 'US', a viralproduct_id) - Aggregations on popular keys
- Exploding arrays unevenly
Diagram
Partitions after hash(key): P0 #### P1 # P2 # P3 ############################ <-- straggler P4 ##
Detection
- Spark UI: task duration max >> median; huge shuffle read on one task
- Pre-check key frequencies:
from pyspark.sql import functions as F
df.groupBy("user_id").count().orderBy(F.desc("count")).show(20)Impact
Long stages, timeouts, disk spill, executor OOM, wasted cluster spend.
Interview tip
Distinguish partition count skew (uneven file sizes) from key skew (hash collisions on hot keys). Fixes differ slightly but both create stragglers.