Apache Iceberg organizes table metadata as a hierarchical tree structure that completely decouples query planning from physical directory listings on cloud storage. At the top of the hierarchy, a catalog maintains an atomic pointer to the current metadata JSON file. Each metadata file defines the table schema, partition specs, and snapshot history, pointing to a manifest list, which in turn points to individual manifest files that catalog individual data files and delete files alongside detailed per-column statistics.
Hierarchical metadata pointers
Each layer in the Iceberg metadata hierarchy serves an explicit role:
- Catalog: Holds the current metadata location and performs atomic compare-and-swap commits when new snapshots are added.
- metadata.json: Tracks table configuration, snapshot history, and schema evolution. When a query targets a snapshot, it reads that snapshot's root manifest list pointer.
- Manifest list: An Avro file containing a list of manifest files. Each entry in the manifest list includes partition summary ranges for the files in that manifest, allowing engines to prune entire manifest files without opening them.
- Manifest file: An Avro file tracking individual data files and positional or equality delete files. Each entry stores the data file path, partition tuple, record count, and lower/upper bounds for every column.
Manifest list and manifest file pruning
Query engines prune data at two distinct metadata stages:
- First, the engine reads the manifest list and inspects partition boundary summaries. Manifests whose partitions cannot match the query filter are pruned immediately.
- Second, the engine opens surviving manifests to inspect per-column min/max bounds, eliminating unneeded data files before generating scan tasks.
Avoiding expensive filesystem listings
In legacy Hive tables, querying a table with 500,000 files across 2,000 partitions required issuing recursive LIST calls to S3, which took minutes and hit API rate limits. In Iceberg, an engine plans queries by reading only compact Avro metadata files, so planning time remains constant even when tables scale to millions of files across petabytes of storage.