The small-files problem is when a dataset is stored as a huge number of tiny files. Spark (and object stores) then spend more time on listing, opening, and scheduling than on reading bytes.
Causes
- Writing with too many output partitions (
repartition(10000)on small data) - Streaming micro-batches each writing many files
- Over-partitioning Hive partitions (partition by high-cardinality columns)
- Frequent appends without compaction
Diagram
Good: 20 x 512MB Parquet files Bad: 50,000 x 8KB Parquet files -> driver listing + task overhead explode
Mitigations
# Reduce output writers before write
df.coalesce(20).write.mode("overwrite").parquet(path)
# Or repartition to a sensible target size
df.repartition(50).write.partitionBy("dt").parquet(path)
# Compact later (lakehouse OPTIMIZE / rewrite jobs)Also: raise maxRecordsPerFile carefully, use AQE coalesce, schedule compaction, prefer fewer partitions for small daily volumes.
Interview tip
Connect small files to cloud listing costs, NameNode pressure (HDFS), slow job startup, and why table formats add compaction utilities.