Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is AQE?

PySpark · Joins & Performance

What is AQE?

Hardpyspark-25
aqeoptimizerspark3performance

Question

What is Adaptive Query Execution (AQE) in Spark?

Solution

Adaptive Query Execution (AQE) lets Spark re-optimize the physical plan at runtime using measured shuffle statistics, instead of relying only on compile-time estimates.

What AQE can do

1. Coalesce shuffle partitions (merge tiny partitions after shuffle) 2. Switch join strategy (e.g., sort-merge -> broadcast when one side is small) 3. Optimize skew joins (split skewed partitions)

Build initial plan
   -> run shuffle / gather stats
   -> adapt: coalesce / change join / skew handling
   -> continue execution

Enable (Spark 3+)

spark.conf.set("spark.sql.adaptive.enabled", "true")
spark.conf.set("spark.sql.adaptive.coalescePartitions.enabled", "true")
spark.conf.set("spark.sql.adaptive.skewJoin.enabled", "true")

Why that matters

Static spark.sql.shuffle.partitions=200 is often wrong for both tiny and huge jobs. AQE reduces small-partition overhead and can fix some skew automatically.

Interview tip

AQE is not magic: bad data models, huge UDFs, and unnecessary wide transformations still dominate. Use AQE plus good file layout and join design.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext