MapReduce is a programming model for large batch jobs on a cluster. You write two main steps:
1. Map: turn each input record into intermediate key/value pairs (in parallel) 2. Reduce: group by key and aggregate/combine values for that key
The framework handles splitting files, shipping data, sorting by key, and restarting failed tasks.
Input splits
| map | map | map |
\ | /
shuffle + sort by key
/ | \
| reduce | reduce |
\ | /
outputClassic word-count mental model
- Map:
(word, 1)for each token - Shuffle: all counts for
hellogo to the same reducer - Reduce: sum the 1s
Interview tip: MapReduce taught the industry "scale by partitioning." Spark kept the idea but kept data in memory across steps to avoid writing every stage to disk.