Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is a DAG? What is Catalyst?

PySpark · Core Concepts

What is a DAG? What is Catalyst?

Mediumpyspark-07
dagcatalystoptimizerexecution

Question

In Spark, what is a DAG, and what is the Catalyst optimizer?

Solution

DAG (Directed Acyclic Graph): when you chain transformations, Spark builds a DAG of stages and tasks. Nodes are operations; edges are dependencies. "Acyclic" means lineage flows one way so Spark can recompute lost partitions.

scan parquet --> filter --> project --> exchange(shuffle) --> aggregate --> write
                 stage 1                         stage 2

Stages split at shuffle boundaries (wide dependencies).

Catalyst

Catalyst is Spark SQL's query optimizer. For DataFrame/Dataset/SQL plans it:

1. Analyzes the logical plan (resolve columns, types) 2. Applies rule-based optimizations (predicate pushdown, constant folding, projection pruning) 3. Builds a cost-based physical plan (join strategy selection when stats exist) 4. Generates Java bytecode (Tungsten / whole-stage codegen) for many operators

SQL / DataFrame API
        |
   Unresolved Logical Plan
        |
   Analyzed Logical Plan
        |
   Optimized Logical Plan   <--- Catalyst rules
        |
   Physical Plan(s)
        |
   Selected Physical Plan
        |
   Codegen + Execution

See the plan in code

df = (
    spark.read.parquet("/data/orders")
    .filter("amount > 100")
    .select("order_id", "amount")
)
df.explain(True)          # parsed / analyzed / optimized / physical
df.explain("formatted")

What to say

DataFrames get Catalyst optimizations. Raw RDD map pipelines usually do not get the same SQL-style plan rewrites.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext