Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is a shuffle?

PySpark · Core Concepts

What is a shuffle?

Mediumpyspark-09
shuffleperformancewide-transformation

Question

What is a shuffle in Spark, and when does it happen?

Solution

A shuffle is the redistribution of data across partitions so that rows that belong together (same key, same range) land on the same executor partition. It is the expensive part of wide transformations.

What happens under the hood

Map side (write shuffle files)     Reduce side (read + compute)
------------------------------     ----------------------------
Executor tasks hash/sort keys  --> network/disk fetch --> tasks aggregate/join

1. Map tasks write shuffle blocks partitioned by key hash (or range). 2. Data moves over the network (and often spills to disk). 3. Reduce tasks read their assigned blocks and continue the plan.

Operations that shuffle

groupBy/agg, most joins, distinct, repartition(n), orderBy, window functions with partitioning, cogroup, etc.

Cost drivers

  • Data volume and skew
  • Number of shuffle partitions (spark.sql.shuffle.partitions, AQE coalescing)
  • Serializer, compression, disk spill
  • Network bandwidth

Example

# Shuffle on user_id
revenue = orders.groupBy("user_id").sum("amount")

# Often avoid shuffle with broadcast if one side is small
from pyspark.sql.functions import broadcast
joined = orders.join(broadcast(dim_users), "user_id")

Interview closer

"Shuffle is necessary for correctness when keys must meet. The craft is cutting shuffled bytes and handling skew."

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext