Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is the small-files problem?

PySpark · Caching & Reliability

What is the small-files problem?

Hardpyspark-29
small-filesiolakehouseperformance

Question

What is the small-files problem in Spark and data lakes?

Solution

The small-files problem is when a dataset is stored as a huge number of tiny files. Spark (and object stores) then spend more time on listing, opening, and scheduling than on reading bytes.

Causes

  • Writing with too many output partitions (repartition(10000) on small data)
  • Streaming micro-batches each writing many files
  • Over-partitioning Hive partitions (partition by high-cardinality columns)
  • Frequent appends without compaction

Diagram

Good:  20 x 512MB Parquet files
Bad:   50,000 x 8KB Parquet files  -> driver listing + task overhead explode

Mitigations

# Reduce output writers before write
df.coalesce(20).write.mode("overwrite").parquet(path)

# Or repartition to a sensible target size
df.repartition(50).write.partitionBy("dt").parquet(path)

# Compact later (lakehouse OPTIMIZE / rewrite jobs)

Also: raise maxRecordsPerFile carefully, use AQE coalesce, schedule compaction, prefer fewer partitions for small daily volumes.

Interview tip

Connect small files to cloud listing costs, NameNode pressure (HDFS), slow job startup, and why table formats add compaction utilities.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext