Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. How do you optimize a slow PySpark job?

PySpark · Caching & Reliability

How do you optimize a slow PySpark job?

Hardpyspark-30
optimizationperformancespark-uibest-practices

Question

How do you approach optimizing a slow PySpark job?

Solution

Use a measurement-first checklist. Do not thrash random configs.

Playbook

1. Reproduce & open Spark UI / SQL plan
2. Find the longest stage (usually shuffle/join/agg)
3. Check input size, row counts, partition counts, skew
4. Fix data layout & query shape
5. Then tune configs / cluster size

Concrete levers

# 1) Filter early + project needed columns
df = spark.read.parquet(path).select("user_id", "amount", "dt").filter("dt = '2024-01-01'")

# 2) Prefer built-in functions over Python UDFs
# 3) Broadcast small dimensions
# 4) Handle skew (AQE skew join, salt, isolate hot keys)
# 5) Fix small files / file sizing on read and write
# 6) Cache only for multi-action reuse
# 7) Enable AQE
spark.conf.set("spark.sql.adaptive.enabled", "true")

Common root causes

  • Unexpected shuffle / Cartesian blow-up
  • Data skew stragglers
  • Reading too much (no partition pruning / column pruning)
  • Python UDFs
  • Tiny files or huge gzip splits that do not split well
  • Driver OOM from collect

Interview structure

"I start from the plan and UI metrics, reduce shuffled bytes, fix skew and file layout, avoid UDFs, then scale executors if the job is legitimately large."

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext