Use a measurement-first checklist. Do not thrash random configs.
Playbook
1. Reproduce & open Spark UI / SQL plan 2. Find the longest stage (usually shuffle/join/agg) 3. Check input size, row counts, partition counts, skew 4. Fix data layout & query shape 5. Then tune configs / cluster size
Concrete levers
# 1) Filter early + project needed columns
df = spark.read.parquet(path).select("user_id", "amount", "dt").filter("dt = '2024-01-01'")
# 2) Prefer built-in functions over Python UDFs
# 3) Broadcast small dimensions
# 4) Handle skew (AQE skew join, salt, isolate hot keys)
# 5) Fix small files / file sizing on read and write
# 6) Cache only for multi-action reuse
# 7) Enable AQE
spark.conf.set("spark.sql.adaptive.enabled", "true")Common root causes
- Unexpected shuffle / Cartesian blow-up
- Data skew stragglers
- Reading too much (no partition pruning / column pruning)
- Python UDFs
- Tiny files or huge gzip splits that do not split well
- Driver OOM from
collect
Interview structure
"I start from the plan and UI metrics, reduce shuffled bytes, fix skew and file layout, avoid UDFs, then scale executors if the job is legitimately large."