Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Open table formats and warehouse lock-in

Snowflake, BigQuery & Databricks · Choosing and Comparing

Open table formats and warehouse lock-in

Mediumwarehouses-56
icebergopen-table-formatslock-ininteroperabilitycatalogs

Question

Why are Snowflake, BigQuery and Databricks all adding Iceberg support?

Solution

Companies want one copy of their data that many engines can read, and they do not want to be locked into one vendor's storage format. Open table formats, especially Apache Iceberg, make that possible. That is why Snowflake, BigQuery and Databricks all added support for it.

The problem they address

Historically, if you loaded data into a warehouse, it was stored in that vendor's proprietary format. Moving to another system meant exporting and reloading. If a team wanted to use Spark on the same data, or another warehouse, they had to copy it. That meant extra cost, lag between copies, and several versions of the truth.

What an open table format does

Iceberg (and Delta Lake and Hudi) add a metadata layer over Parquet files in object storage. It records the schema, partitions, snapshots and which files make up the table, and it supports transactions, schema changes and time travel. Any engine that understands the format can read the same files, and the table's history, safely.

Parquet files in your bucket + Iceberg metadata
   ^           ^            ^            ^
Snowflake   BigQuery     Databricks   Spark / Trino / Flink

Catalogs make it practical

Engines need to find tables and agree on their latest state. An Iceberg REST catalog (such as Polaris, now Apache Polaris, Unity Catalog's Iceberg endpoint, or AWS Glue) provides a shared place for this. With a common catalog API, several engines can use the same tables.

Why each vendor joined in

Customers asked for it, and it lowers the barrier to adopting the platform, since data can stay where it is. It also shifts competition to the quality of the engine, governance and features, instead of format lock-in.

Trade-offs

  • Some native features and performance optimizations work best on the vendor's own tables, so Iceberg tables may lag in features.
  • Governance is split: access to the files in the bucket and access through the catalog must be consistent.
  • Someone has to run maintenance: compaction, snapshot expiry, and orphan file cleanup, unless the platform does it.
  • Multiple writers need careful coordination, since commits go through the catalog.

A balanced view

Open formats reduce lock-in, but they do not remove it. Your catalog, governance, pipelines and skills still tie you to some platform. A sensible approach: keep shared, long-lived datasets in open formats, and use native tables where speed or features matter most.

PreviousNext