Data skew means a few keys hold a disproportionate share of the data. Tasks on those keys become stragglers while other tasks finish and sit idle.
join / groupBy key = country US: ############################ (huge) IN: ########## DE: ## IS: # One executor stuck on US; cluster looks "busy" but progress crawls.
Where skew shows up
- Joins and aggregations on popular keys (
null,unknown, big customers) - Hot partitions in Kafka
- Uneven file sizes after bad partitioning
Mitigations
1. Filter / isolate hot keys and process them separately 2. Salting: add a random suffix to hot keys, then aggregate twice 3. Broadcast join the small side instead of shuffling both 4. Adaptive skew join (Spark AQE) when available 5. Repartition with a better key or more partitions
Interview tip: Define skew as uneven key distribution causing stragglers, then name salting or broadcast as fixes. Mention checking Spark stage timelines for one long task.