They both involve "partitions," but at different layers.
repartition(...) (DataFrame API)
Controls how many in-memory partitions (and optionally by which keys) Spark uses for execution of the next stages. It triggers a shuffle when redistributing.
partitionBy(...) (writer API)
Controls directory layout on disk when writing a table/files (Hive-style partitions).
# Execution partitioning (shuffle now)
df2 = df.repartition(100, "user_id")
# Storage partitioning (on write)
df2.write.mode("overwrite").partitionBy("event_date").parquet("/lake/events")Diagram
repartition(user_id) -> Spark partitions for compute partitionBy(event_date) -> /event_date=2024-01-01/files...
Interview tip
You often repartition to control file count / balance, and partitionBy so readers can prune by date/region. Do not partition storage by an ultra-high cardinality key.