Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. How do you fix data skew?

PySpark · Joins & Performance

How do you fix data skew?

Hardpyspark-24
skewsaltingaqeperformance

Question

How do you fix or mitigate data skew in PySpark jobs?

Solution

Mitigation depends on whether skew is on join keys, aggregations, or input file sizes.

Techniques

1. Filter / isolate hot keys: process the mega-key separately, union results. 2. Salting: add a random salt to hot keys on both sides (for joins) or on the group key, aggregate partial then final. 3. AQE skew join optimization: enable adaptive skew join so Spark splits skewed partitions. 4. Broadcast the small side so the large skewed fact is not shuffle-joined. 5. Increase parallelism cautiously (more shuffle partitions) - helps some imbalance, not true hot keys. 6. Repartition by better keys or two-phase aggregation (reduceByKey-style partial aggs via DataFrame groupBy already does partial aggregation).

Salting sketch

from pyspark.sql import functions as F

salt_n = 16
hot = fact.withColumn("salt", (F.rand() * salt_n).cast("int"))
dim_exp = dim.withColumn("salt", F.explode(F.array([F.lit(i) for i in range(salt_n)])))

joined = hot.join(dim_exp, ["user_id", "salt"])

Diagram

Before: key=K ########################  -> 1 task
After salt: K#0 ####  K#1 #### ... K#15 #### -> many tasks

Interview tip

Lead with measurement (UI + key counts), then pick the cheapest fix (broadcast / AQE) before hand-rolled salting.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext