Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. maxRecordsPerFile and controlling output files

PySpark · Joins & Data Layout

maxRecordsPerFile and controlling output files

Mediumpyspark-71
output-filesrepartitionsmall-filespartitionby

Question

How do you control how many output files Spark writes?

Solution

The number of files Spark writes is, roughly, the number of write tasks multiplied by the number of partition directories each task has data for. That is why a job with 200 tasks writing to 365 date partitions can create tens of thousands of files.

How the count builds up

Say the DataFrame has 200 partitions (after a shuffle) and you write with partitionBy("order_date"). Each task holds rows from many dates, and writes one file per date it holds. If each of the 200 tasks has rows for all 30 days in the month, you get 200 times 30 = 6,000 files for 30 partitions. About 200 tiny files per day.

Get one file per partition directory

Repartition by the same column before writing, so that all rows for a date land in the same task:

(df.repartition("order_date")
   .write.partitionBy("order_date")
   .mode("overwrite")
   .parquet("/lake/orders"))

Now each date is written by one task and produces one file. If one date is huge, that single file may be too large and the one task becomes a bottleneck. To get a few files per day, add a second column or a random number: repartition(8, "order_date", F.rand()), or use repartitionByRange.

Capping file size

maxRecordsPerFile splits a file when it reaches a row count:

df.write.option("maxRecordsPerFile", 2_000_000).partitionBy("order_date").parquet(path)

Use it as a safety net so that no file grows beyond a reasonable size, say a few hundred MB.

coalesce versus repartition

coalesce(n) reduces partitions without a full shuffle, so it is cheap, but it can leave uneven partitions and can reduce the parallelism of earlier steps in the same stage. repartition(n) does a full shuffle and gives even partitions. Use coalesce for a modest reduction at the end, repartition when you need balance or a column-based layout.

Why it matters

Thousands of small files make later reads slow (listing and opening cost more than reading), and put load on the storage metadata service. Aim for files of roughly 128 MB to 1 GB. A scheduled compaction job, or the table format's OPTIMIZE, can fix files that were already written small. On Databricks, optimized writes and auto compaction can do some of this for you.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext