Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. repartition() vs partitionBy()

PySpark · Partitioning & Pruning

repartition() vs partitionBy()

Mediumpyspark-36
repartitionpartitionbypartitions

Question

What is the difference between repartition() and partitionBy() in PySpark?

Solution

They both involve "partitions," but at different layers.

repartition(...) (DataFrame API)

Controls how many in-memory partitions (and optionally by which keys) Spark uses for execution of the next stages. It triggers a shuffle when redistributing.

partitionBy(...) (writer API)

Controls directory layout on disk when writing a table/files (Hive-style partitions).

# Execution partitioning (shuffle now)
df2 = df.repartition(100, "user_id")

# Storage partitioning (on write)
df2.write.mode("overwrite").partitionBy("event_date").parquet("/lake/events")

Diagram

repartition(user_id)     -> Spark partitions for compute
partitionBy(event_date)  -> /event_date=2024-01-01/files...

Interview tip

You often repartition to control file count / balance, and partitionBy so readers can prune by date/region. Do not partition storage by an ultra-high cardinality key.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext