Data teams must regularly expire old table snapshots and purge orphan files to prevent abandoned data files from driving cloud storage bills exponentially higher. Table formats preserve older data files to support point-in-time time travel queries and concurrent read isolation, but these unreferenced files remain on object storage forever unless explicitly deleted by maintenance commands. However, running expiration commands with an aggressive retention threshold risks terminating active long-running queries or corrupting in-flight write operations.
The balance between time travel and storage growth
Every time an UPDATE, DELETE, MERGE, or compaction runs in Delta Lake or Iceberg, modified Parquet files are logically superseded by new files. Because the table maintains historical snapshots in its transaction log or manifest tree:
- Older snapshots keep referencing the superseded Parquet files so analysts can run historical queries.
- In high-throughput ingestion pipelines, obsolete files quickly consume 70 to 90 percent of total storage volume if snapshot cleanup is neglected.
- The expire_snapshots procedure in Iceberg or VACUUM in Delta removes commit metadata older than a specified retention window and deletes the physical files that are no longer referenced by any retained snapshot.
Why orphan files appear
Orphan files represent a different storage leak caused by failed write tasks:
- If a Spark executor crashes while writing speculative Parquet files to S3, or if a job throws an out-of-memory exception before committing, those written data files are never registered in the transaction log.
- Standard snapshot expiration will never delete them because the transaction log does not know they exist.
- Teams must run dedicated orphan cleanup routines (such as Iceberg's remove_orphan_files) that scan the physical storage directory, compare existing objects against all known manifest files, and delete untracked objects.
Setting safe retention intervals
Managing retention requires balancing safety against cost:
- Setting retention too short (such as 1 hour) creates critical failures: if an analytical query takes 90 minutes to run, a concurrent vacuum can delete the files underneath it, triggering FileNotFoundException crashes.
- Never set retention to zero on active tables because speculative files from currently executing jobs will be mistakenly wiped as orphans.
- A battle-tested default is seven days for snapshot retention, scheduled as an automated weekly or daily maintenance DAG.