A stage is a set of tasks that can run together without a shuffle boundary. Spark splits a job's DAG into stages at wide dependencies.
Hierarchy
Application
-> Jobs (one per action)
-> Stages (pipelined narrow transforms)
-> Tasks (one per partition)Diagram
Stage 0: scan -> filter -> project
\
shuffle exchange
/
Stage 1: aggregate -> project -> writeWhy that matters
- The Spark UI's longest stage is usually where you optimize first
- Task count in a stage ~= number of partitions for that stage
- Stragglers inside a stage often mean skew or uneven file sizes
# Each action triggers a job with one or more stages
df.filter("amount > 0").groupBy("user_id").sum("amount").write.parquet("/out")
# typically: stage for map-side prep + stage for post-shuffle aggregation + writeInterview tip
"Jobs are triggered by actions; stages are separated by shuffles; tasks run per partition." That sentence alone answers a lot of Spark interview questions.