Small files means a table or path is made of thousands/millions of tiny files (KB–few MB) instead of fewer larger columnar files. Listing, opening, and planning over them kills job time and stresses the metastore/object store.
How they appear
- Streaming sinks writing every micro-batch
repartitiontoo high before write- Many parallel tasks each writing one file
- Over-partitioned Hive-style paths (
date=.../hour=.../minute=...)
Bad: day/ f1(20KB) f2(20KB) ... f50000(20KB) Good: day/ part-00001.parquet (256MB) part-00002.parquet (256MB)
Fixes
- Coalesce/repartition before write to target file size (~128–512MB Parquet often)
- Compaction jobs (Delta OPTIMIZE, Iceberg rewrite_data_files)
- Fewer partitions; batch micro-batch outputs
- Auto-compact features in lakehouse formats
Interview tip: Link cause (too many writers/partitions) → symptom (slow LIST + task startup) → fix (compaction / coalesce).