Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. File compaction

File formats & storage · Extra High-Value

File compaction

Mediumformat-17
compactionOPTIMIZEsmall filesHudiIcebergDelta

Question

What is file compaction in a data lake or lakehouse?

Solution

File compaction rewrites many small (or fragmented) data files into fewer, larger, well-sized files, and updates table metadata so readers use the new set.

Why you need it

Streaming and frequent micro-batch appends create thousands of tiny files. Deletes/updates in MoR-style tables leave log fragments. Compaction restores scan efficiency and healthy file sizes.

Before: 10,000 x 2MB files in dt=2024-01-02/
After:     40 x 500MB files + metadata commit removing old files

How it shows up by format

  • Delta: OPTIMIZE (plus vacuum for old files)
  • Iceberg: rewrite data files / expire snapshots
  • Hudi: compaction especially important for Merge-on-Read
  • DIY lakes: Spark job that reads partition and overwrites with fewer files

Operational notes

  • Compaction costs compute; schedule it (nightly / after heavy ingest)
  • Coordinate with retention/vacuum so readers are not broken mid-cleanup
  • Aim for target file sizes your engine likes (often ~128MB–1GB range, stack-dependent)

Interview tip: Compaction is the cure for the small-files problem and for MoR read amplification. Mention metadata commit, not only "merge files."

PreviousNext