The Spark UI is Spark's web dashboard for an application (typically port 4040 on the driver, or via YARN/History Server for finished apps).
Key tabs
Jobs -> Stages -> Tasks | | | +-- task duration, spill, GC, shuffle read/write +-- DAG visualization, event timeline SQL / DataFrame - query plans, scanned bytes, spill metrics Storage - cached RDDs/DataFrames Environment / Executors - configs, memory, cores, failed executors
What I check first on a slow job
1. SQL tab: physical plan, whether broadcast was used, partition filters. 2. Stages: longest stage, whether it is a shuffle stage. 3. Tasks: skewed max/median duration (data skew signal). 4. Shuffle read/write and spill (memory/disk). 5. GC time and executor lost / speculation.
Practical workflow
slow action -> find Job -> find longest Stage -> inspect straggler tasks -> relate back to join/agg key and input size
Tip in code / local runs
# UI usually at http://localhost:4040 while the app is alive
spark = SparkSession.builder.appName("debug-me").getOrCreate()
df.groupBy("user_id").count().write.mode("overwrite").parquet("/tmp/out")
# Inspect Jobs/Stages/SQL tabs while this action runs (or use History Server after)Interview tip
Mention History Server for post-mortem of production jobs. Reading the plan beats guessing config knobs.