Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is a Spark stage?

PySpark · Partitioning & Pruning

What is a Spark stage?

Mediumpyspark-40
stagedagexecutionspark-ui

Question

What is a Spark stage?

Solution

A stage is a set of tasks that can run together without a shuffle boundary. Spark splits a job's DAG into stages at wide dependencies.

Hierarchy

Application
  -> Jobs          (one per action)
     -> Stages     (pipelined narrow transforms)
        -> Tasks   (one per partition)

Diagram

Stage 0: scan -> filter -> project
                 \
                  shuffle exchange
                 /
Stage 1: aggregate -> project -> write

Why that matters

  • The Spark UI's longest stage is usually where you optimize first
  • Task count in a stage ~= number of partitions for that stage
  • Stragglers inside a stage often mean skew or uneven file sizes
# Each action triggers a job with one or more stages
df.filter("amount > 0").groupBy("user_id").sum("amount").write.parquet("/out")
# typically: stage for map-side prep + stage for post-shuffle aggregation + write

Interview tip

"Jobs are triggered by actions; stages are separated by shuffles; tasks run per partition." That sentence alone answers a lot of Spark interview questions.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext