An Apache Iceberg catalog is the central authority that maintains the pointer to the current metadata JSON file for every table and coordinates atomic commits during transactions. Available catalog implementations include AWS Glue, Hive Metastore, Nessie, JDBC, and the open REST catalog specification. The REST catalog specification matters because it standardizes an open HTTP protocol that decouples query engines from specific catalog backend implementations, allowing any compute engine to interact with modern catalogs like Apache Polaris, Databricks Unity Catalog, and Snowflake Open Catalog.
Atomic state coordination via catalogs
In an Iceberg table, data and metadata files sit on object storage, but an engine cannot discover the active state without the catalog:
- When a reader queries a table, it asks the catalog for the URI of the current metadata file.
- When a writer commits a new snapshot, it writes the new metadata JSON and instructs the catalog to perform an atomic compare-and-swap operation, replacing the old metadata pointer with the new one.
- Historically, catalogs were tied to specific technologies: Hive Metastore used Thrift, Glue used AWS SDKs, and JDBC required direct database connections. Each engine needed custom client libraries and credentials to speak each protocol.
The universal REST specification
The Iceberg REST catalog specification changes this by defining a standardized OpenAPI REST interface:
- Any compute engine implementing the REST client (such as Spark, Trino, Flink, DuckDB, or StarRocks) can query any catalog implementing the spec without proprietary plugins.
- Centralized governance systems such as Apache Polaris, Databricks Unity Catalog, Snowflake Open Catalog, and AWS S3 Tables expose REST endpoints natively.
- Authentication, role-based access control, and credential vending (such as passing short-lived S3 token credentials down to compute engines) are handled through standard OAuth2 HTTP headers.
Interoperable multi-engine architecture
This specification enables a true multi-engine lakehouse. A single dataset can be ingested by Flink, transformed in Spark, and queried simultaneously by Trino and Snowflake, all authenticating against a single REST catalog that enforces consistent access rules and transactional commits across the enterprise.