Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Hadoop MapReduce vs Spark

Batch & Streaming · Batch Processing

Hadoop MapReduce vs Spark

Easystream-12
SparkHadoopMapReduceperformance

Question

How does Apache Spark differ from classic Hadoop MapReduce?

Solution

Hadoop MapReduce runs multi-step jobs as a chain of map→shuffle→reduce stages that typically write intermediate results to disk (HDFS). Great for huge, simple batch; slow for iterative or multi-stage pipelines.

Spark builds a DAG of stages and keeps intermediate data in memory when it can. One Spark app can do many maps/joins/filters without a full HDFS round-trip each time.

MapReduce-style multi-step
  step1 -> HDFS -> step2 -> HDFS -> step3

Spark DAG
  read -> filter -> join -> agg -> write
          (memory / local spill between stages)

Other Spark advantages interviewers expect

  • Richer APIs: SQL, DataFrames, MLlib, streaming
  • Lazy evaluation + Catalyst/Tungsten optimizations
  • Interactive and iterative workloads (ML, graph)

Interview tip: "MR = disk-heavy stage machine. Spark = in-memory DAG engine with SQL." Mention both still shuffle across the network for wide dependencies.

PreviousNext