Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Narrow vs wide transformations

Batch & Streaming · Batch Processing

Narrow vs wide transformations

Mediumstream-16
narrowwideshufflestages

Question

What are narrow and wide transformations in Spark?

Solution

Narrow transformations: each output partition depends on one input partition. No shuffle. Examples: map, filter, union (often), column projections.

Wide transformations: output partitions depend on many input partitions. Data must be reshuffled across the network. Examples: groupByKey, reduceByKey, join (typical), distinct, repartition.

Narrow                         Wide (shuffle boundary)
P0 -> P0'                      P0 --\
P1 -> P1'                      P1 ---+--> exchange by key --> new partitions
P2 -> P2'                      P2 --/

Wide deps create stage boundaries. More wide deps → more stages → usually more cost.

Interview tip: "Wide = shuffle = stage break." Prefer narrow pipes when you can; when you must shuffle, watch skew and partition count.

PreviousNext