Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. The small files problem

File formats & storage · Performance & Layout

The small files problem

Mediumformat-11
small filescompactionSparkS3performance

Question

What is the small-files problem in data lakes?

Solution

The small-files problem is when a dataset is stored as a huge number of tiny files. Engines and object stores then spend more time on listing, opening, planning, and task scheduling than on reading useful bytes.

Common causes

  • Writing with far too many output partitions (repartition(10000) on modest data)
  • Streaming micro-batches each flushing many files
  • Over-partitioning (partition by hour *and* high-cardinality columns)
  • Frequent appends without compaction
Good:  20 x ~512MB Parquet files
Bad:   50,000 x 8KB Parquet files
       → slow LIST on S3, huge Spark job startup, NameNode pressure on HDFS

Mitigations

  • coalesce / repartition to a target file count or size before write
  • Raise maxRecordsPerFile thoughtfully
  • Schedule compaction (Delta OPTIMIZE, Iceberg rewrite, Hudi compact)
  • Prefer coarser partitions for small daily volumes

Interview tip: Connect small files to cloud LIST costs, slow job starts, and why lakehouse formats ship compaction utilities. Target roughly hundreds-of-MB files unless your engine docs say otherwise.

PreviousNext