DAG (Directed Acyclic Graph): when you chain transformations, Spark builds a DAG of stages and tasks. Nodes are operations; edges are dependencies. "Acyclic" means lineage flows one way so Spark can recompute lost partitions.
scan parquet --> filter --> project --> exchange(shuffle) --> aggregate --> write
stage 1 stage 2Stages split at shuffle boundaries (wide dependencies).
Catalyst
Catalyst is Spark SQL's query optimizer. For DataFrame/Dataset/SQL plans it:
1. Analyzes the logical plan (resolve columns, types) 2. Applies rule-based optimizations (predicate pushdown, constant folding, projection pruning) 3. Builds a cost-based physical plan (join strategy selection when stats exist) 4. Generates Java bytecode (Tungsten / whole-stage codegen) for many operators
SQL / DataFrame API
|
Unresolved Logical Plan
|
Analyzed Logical Plan
|
Optimized Logical Plan <--- Catalyst rules
|
Physical Plan(s)
|
Selected Physical Plan
|
Codegen + ExecutionSee the plan in code
df = (
spark.read.parquet("/data/orders")
.filter("amount > 100")
.select("order_id", "amount")
)
df.explain(True) # parsed / analyzed / optimized / physical
df.explain("formatted")What to say
DataFrames get Catalyst optimizations. Raw RDD map pipelines usually do not get the same SQL-style plan rewrites.