Table formats like Delta Lake, Apache Iceberg, and Apache Hudi achieve ACID transactions on cloud object storage by introducing a decoupled metadata transaction log that tracks exactly which physical Parquet files constitute valid table snapshots. Instead of modifying files in place or relying on object storage directory listings, writers produce immutable data files and then execute an atomic commit that appends a new state to the transaction log. Readers query the table by referencing a specific committed log version, providing isolated, consistent snapshot reads while background maintenance cleans up stale files.
Decoupling data files from committed state
In a traditional Hive metastore setup, a table was defined merely as a folder path on storage. If a write job failed halfway through, half-written files remained visible, producing dirty reads and corrupted aggregations. Table formats eliminate folder-based state:
- Atomicity: When an insert or merge executes, new Parquet files are written to storage. Only after all writes succeed does the writer attempt to commit the new snapshot metadata. If the commit fails, the table remains untouched from the perspective of readers.
- Consistency and Isolation: Readers query a specific snapshot point in time. A reader starting a query at version 10 will never see uncommitted files or rows added by concurrent writes in version 11.
- Durability: Once the transaction log entry is committed to durable object storage or a catalog, the new state is permanent.
Atomicity and optimistic concurrency
Concurrency is managed through optimistic concurrency control. When two jobs write concurrently, both read the current snapshot (say version 5) and write candidate files:
- The first writer successfully commits version 6.
- The second writer then detects that version 5 is outdated, checks whether its modified partitions or files conflict with version 6, and either replays its commit or throws a write conflict exception.
- This optimistic check prevents overwriting data without requiring heavy distributed locks.
Garbage collection and retention
Because table formats never delete old Parquet files synchronously during writes, previous file generations accumulate on disk to support historical time travel queries. Left unchecked, storage costs explode. Teams run scheduled maintenance routines, such as VACUUM in Delta or expire_snapshots in Iceberg, to physically purge data files no longer referenced by any active snapshot past a configured retention threshold.