Lazy evaluation means Spark does not run transformations when you call them. It records a plan (lineage) and executes only when an action forces a result.
Why that helps
Spark can optimize the full pipeline before running: combine filters, prune columns, push predicates into file readers, and avoid unnecessary shuffles.
Flow
read -> filter -> select -> join -> groupBy (transformations: build plan)
|
action (count/show/write)
|
execute optimized jobsExample
df = spark.read.parquet("/data/orders") # lazy
f = df.filter("amount > 100") # lazy
s = f.select("order_id", "amount") # still lazy
print(s.count()) # action: job runs now
s.write.mode("overwrite").parquet("/out/orders") # action: job runs againInterview pitfalls
- Calling many actions (
count,show,collect) recomputes from scratch unless you cache/persist or write intermediate results. - Errors in transformations often appear only at action time, which can surprise beginners.
One-liner
"Transformations describe work; actions trigger work."