The small-files problem is when a dataset is stored as a huge number of tiny files. Engines and object stores then spend more time on listing, opening, planning, and task scheduling than on reading useful bytes.
Common causes
- Writing with far too many output partitions (
repartition(10000)on modest data) - Streaming micro-batches each flushing many files
- Over-partitioning (partition by hour *and* high-cardinality columns)
- Frequent appends without compaction
Good: 20 x ~512MB Parquet files
Bad: 50,000 x 8KB Parquet files
→ slow LIST on S3, huge Spark job startup, NameNode pressure on HDFSMitigations
coalesce/repartitionto a target file count or size before write- Raise
maxRecordsPerFilethoughtfully - Schedule compaction (Delta
OPTIMIZE, Iceberg rewrite, Hudi compact) - Prefer coarser partitions for small daily volumes
Interview tip: Connect small files to cloud LIST costs, slow job starts, and why lakehouse formats ship compaction utilities. Target roughly hundreds-of-MB files unless your engine docs say otherwise.