Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. The small files problem

Batch & Streaming · Batch Processing

The small files problem

Mediumstream-20
small filesParquetcompactionlake

Question

What is the small files problem in data lakes, and how do you fix it?

Solution

Small files means a table or path is made of thousands/millions of tiny files (KB–few MB) instead of fewer larger columnar files. Listing, opening, and planning over them kills job time and stresses the metastore/object store.

How they appear

  • Streaming sinks writing every micro-batch
  • repartition too high before write
  • Many parallel tasks each writing one file
  • Over-partitioned Hive-style paths (date=.../hour=.../minute=...)
Bad:  day/  f1(20KB) f2(20KB) ... f50000(20KB)
Good: day/  part-00001.parquet (256MB) part-00002.parquet (256MB)

Fixes

  • Coalesce/repartition before write to target file size (~128–512MB Parquet often)
  • Compaction jobs (Delta OPTIMIZE, Iceberg rewrite_data_files)
  • Fewer partitions; batch micro-batch outputs
  • Auto-compact features in lakehouse formats

Interview tip: Link cause (too many writers/partitions) → symptom (slow LIST + task startup) → fix (compaction / coalesce).

PreviousNext