Snowflake can use Apache Iceberg tables, where the data and metadata files sit in your own cloud storage in the open Iceberg format, rather than in Snowflake's internal storage. There are two styles, depending on who runs the catalog.
Snowflake-managed Iceberg tables
Snowflake is the Iceberg catalog. You create the table in Snowflake, point it at an external volume (your bucket), and Snowflake reads and writes it, handling the metadata. You get most Snowflake features. Because the files are standard Iceberg, other engines can read the same table, for example Spark or Trino, through a catalog that exposes it.
Externally managed Iceberg tables
Another system owns the catalog, such as AWS Glue or an Iceberg REST catalog like Snowflake's open-source-based Open Catalog (Polaris). Snowflake reads the table through the catalog integration. Spark or Flink may be writing to it, and Snowflake queries it. Write support from Snowflake into externally managed tables is more limited, so check the current documentation.
CREATE ICEBERG TABLE orders_ice EXTERNAL_VOLUME = 'my_s3_volume' CATALOG = 'snowflake' BASE_LOCATION = 'orders/';
Why companies want this
One copy of the data readable by several engines, and no lock-in to one vendor's storage format. A team might ingest with Spark, serve BI from Snowflake, and run an ML job from another engine, all on the same files.
Trade-offs against native tables
- Some Snowflake features work differently or are missing on Iceberg tables, and some performance optimizations are only on native storage.
- You manage the bucket, its permissions and its cost, and you are responsible for file layout. Compaction and snapshot expiry need attention.
- Cross-engine writes need care, since only one writer should own the catalog metadata.
Say that you would keep hot, heavily used tables native for performance and features, and use Iceberg where multi-engine access matters more.