Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Iceberg format v3

File formats & storage · Table Formats in Depth

Iceberg format v3

Mediumfile-formats-33
icebergiceberg-v3table-formatsvariantspecifications

Question

What's new in Iceberg format version 3?

Solution

Apache Iceberg format version 3 (v3) is an upgraded specification that expands the capabilities of the lakehouse metadata layer with native performance and type improvements. Major additions include native deletion vectors for efficient merge-on-read, row lineage identifiers for fine-grained change tracking, the semi-structured VARIANT type, geospatial types, default column values, and nanosecond timestamp precision. While the v3 specification has been finalized by the Apache Iceberg community, production engine support across Spark, Trino, Snowflake, and DuckDB is rolling out gradually through 2025 and 2026.

Core capabilities introduced in v3

Iceberg v3 addresses several operational limitations that existed in the v2 specification:

  • Deletion vectors: In Iceberg v2, merge-on-read relied on positional delete files, which were full Avro files tracking file paths and row positions. Reading many delete files created memory pressure and slow joins. Iceberg v3 replaces them with compact deletion vectors stored as compressed Roaring Bitmaps, aligning performance with Delta Lake.
  • Row lineage and row IDs: v3 introduces immutable row identifiers that persist across updates and compaction rewrites. This enables incremental processing frameworks to track individual records as they move through transformation layers.
  • Default column values: When adding a new column to a table schema, v3 allows defining a default value in metadata. Historical files return this default value without requiring a backfill rewrite.

Modern data types and row tracking

The type system in v3 receives major modernizations:

  • The VARIANT type brings standardized, shreddable semi-structured JSON storage to open lakehouses, allowing query engines to push down predicates into nested JSON subfields without treating them as opaque strings.
  • Nanosecond timestamps accommodate high-frequency trading and IoT sensor data that previously suffered rounding errors under microsecond limits.
  • Geospatial types provide native support for geometries and geography objects, standardizing spatial analytics across engines.

Engine rollout and production considerations

Before adopting Iceberg v3 in your pipelines, audit your query stack. While write engines like Apache Spark lead implementation, external engines like AWS Athena, Trino, or older warehouse connectors may lag in supporting v3 deletion vectors or the VARIANT type. Verify compatibility across all reading engines before flipping the table format version in production catalogs.

PreviousNext