Mitigation depends on whether skew is on join keys, aggregations, or input file sizes.
Techniques
1. Filter / isolate hot keys: process the mega-key separately, union results. 2. Salting: add a random salt to hot keys on both sides (for joins) or on the group key, aggregate partial then final. 3. AQE skew join optimization: enable adaptive skew join so Spark splits skewed partitions. 4. Broadcast the small side so the large skewed fact is not shuffle-joined. 5. Increase parallelism cautiously (more shuffle partitions) - helps some imbalance, not true hot keys. 6. Repartition by better keys or two-phase aggregation (reduceByKey-style partial aggs via DataFrame groupBy already does partial aggregation).
Salting sketch
from pyspark.sql import functions as F
salt_n = 16
hot = fact.withColumn("salt", (F.rand() * salt_n).cast("int"))
dim_exp = dim.withColumn("salt", F.explode(F.array([F.lit(i) for i in range(salt_n)])))
joined = hot.join(dim_exp, ["user_id", "salt"])Diagram
Before: key=K ######################## -> 1 task After salt: K#0 #### K#1 #### ... K#15 #### -> many tasks
Interview tip
Lead with measurement (UI + key counts), then pick the cheapest fix (broadcast / AQE) before hand-rolled salting.