A single lakehouse table can be read across Delta Lake, Apache Iceberg, and Apache Hudi engines simultaneously without duplicating underlying data files by generating multiple metadata formats over a single shared set of Parquet files. Delta Lake Universal Format (UniForm) automatically generates Iceberg metadata (and Hudi metadata) in the background whenever a Delta commit occurs, allowing external Iceberg-native readers like Snowflake or DuckDB to query the Delta table natively. Similarly, open-source projects like Apache XTable provide bi-directional metadata translation across all three major formats.
Single storage layer with multi-format metadata
Historically, choosing a table format locked teams into specific compute engines: Databricks favored Delta Lake, while Snowflake and Trino favored Apache Iceberg. Migrating between formats previously required duplicating petabytes of storage:
- Delta UniForm: When enabled on a Delta table, the Delta writer commits its standard _delta_log entries and then asynchronously generates corresponding Iceberg metadata.json, manifest lists, and manifest files pointing to the exact same Parquet data files.
- Apache XTable: Operates as an independent translation tool. It inspects source metadata (from Hudi, Delta, or Iceberg) and generates synchronized target metadata files, allowing any engine to read the table through its preferred interface without data copies.
Translation via Apache XTable
Apache XTable functions by mapping table abstractions:
- It translates schemas, partition specs, snapshot histories, and file-level statistics between format standards.
- A single pipeline can write to Delta, and an XTable sync job creates Iceberg and Hudi metadata files pointing to the exact same Parquet files in S3.
- This approach avoids copying terabytes of analytical data just to serve different query engines.
Compatibility limitations and ecosystem direction
However, metadata translation comes with operational limitations:
- Translation is restricted to features supported across all target formats. Advanced capabilities like Delta deletion vectors, Iceberg equality deletes, or proprietary catalog extensions cannot always map directly between specifications.
- Metadata synchronization can experience slight latency: if Iceberg metadata generation is asynchronous, an external Iceberg reader might see table updates seconds or minutes after the Delta write finishes.
- The aggressive format competition among lakehouse vendors has cooled down significantly as cloud data platforms converge on Parquet as the common storage bedrock.