File compaction rewrites many small (or fragmented) data files into fewer, larger, well-sized files, and updates table metadata so readers use the new set.
Why you need it
Streaming and frequent micro-batch appends create thousands of tiny files. Deletes/updates in MoR-style tables leave log fragments. Compaction restores scan efficiency and healthy file sizes.
Before: 10,000 x 2MB files in dt=2024-01-02/ After: 40 x 500MB files + metadata commit removing old files
How it shows up by format
- Delta:
OPTIMIZE(plus vacuum for old files) - Iceberg: rewrite data files / expire snapshots
- Hudi: compaction especially important for Merge-on-Read
- DIY lakes: Spark job that reads partition and overwrites with fewer files
Operational notes
- Compaction costs compute; schedule it (nightly / after heavy ingest)
- Coordinate with retention/vacuum so readers are not broken mid-cleanup
- Aim for target file sizes your engine likes (often ~128MB–1GB range, stack-dependent)
Interview tip: Compaction is the cure for the small-files problem and for MoR read amplification. Mention metadata commit, not only "merge files."