Hadoop MapReduce runs multi-step jobs as a chain of map→shuffle→reduce stages that typically write intermediate results to disk (HDFS). Great for huge, simple batch; slow for iterative or multi-stage pipelines.
Spark builds a DAG of stages and keeps intermediate data in memory when it can. One Spark app can do many maps/joins/filters without a full HDFS round-trip each time.
MapReduce-style multi-step
step1 -> HDFS -> step2 -> HDFS -> step3
Spark DAG
read -> filter -> join -> agg -> write
(memory / local spill between stages)Other Spark advantages interviewers expect
- Richer APIs: SQL, DataFrames, MLlib, streaming
- Lazy evaluation + Catalyst/Tungsten optimizations
- Interactive and iterative workloads (ML, graph)
Interview tip: "MR = disk-heavy stage machine. Spark = in-memory DAG engine with SQL." Mention both still shuffle across the network for wide dependencies.