Apache Iceberg is an open table format for huge analytic tables on object storage. Like Delta, it tracks which data files belong to which table snapshot, so readers get ACID, time travel, and safe schema evolution.
How Iceberg thinks
catalog / metastore
|
Iceberg metadata (snapshots, manifests)
|
data files (usually Parquet, sometimes ORC/Avro)Iceberg uses manifests and snapshot metadata so planning can prune files without listing every object in S3 for every query. That matters at massive scale.
Iceberg vs Delta (practical)
| Area | Iceberg | Delta Lake | |-------------------|----------------------------------------------|---------------------------------------------| | Spec | Open community table format | Open format; strong Databricks lineage | | Metadata model | Snapshots + manifests | Transaction log (_delta_log) | | Engine story | Broad multi-engine (Spark, Trino, Flink…) | Excellent Spark/Databricks; growing readers | | Hidden partitioning | Strong "partition evolution" story | Partitioning + liquid clustering / Z-Order | | Upserts | Supported (MERGE, etc. via engines) | Mature MERGE story |
When interviews want a pick
- Multi-engine lake across Spark + Trino + Flink → Iceberg often cited
- Databricks-first / Spark-centric lakehouse → Delta often cited
Interview tip: Both solve "reliable tables on files." Contrast metadata design and ecosystem, not "one has ACID and the other doesn't."