Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Deletion vectors

File formats & storage · Table Formats in Depth

Deletion vectors

Mediumfile-formats-32
deletion-vectorsdelta-lakeicebergperformanceoptimization

Question

What are deletion vectors?

Solution

Deletion vectors are compact bitmap data structures that record which specific row positions in an existing Parquet data file have been marked as deleted without rewriting the underlying file. Adopted in Delta Lake and introduced in Apache Iceberg format version 3 to replace positional delete files, deletion vectors drastically reduce write amplification during DELETE, UPDATE, and MERGE operations. Query engines read the deletion vector alongside the base Parquet file to skip deleted row offsets in memory, deferring physical removal of rows until asynchronous compaction runs.

Roaring bitmaps instead of file duplication

Base Data File: sales_01.parquet (Row indices 0 to 999,999)
  +
Deletion Vector: sales_01.bin (Roaring bitmap: marks row 42, 105, 9021)
  =
Reader Scan: Emits rows 0..999,999 skipping offsets [42, 105, 9021]

Deletion vectors rely on compressed bitset structures, typically Roaring Bitmaps:

  • Instead of writing a separate Avro or Parquet positional delete file that lists 64-bit row offsets, a deletion vector stores compressed bit arrays directly. A single bit set to 1 indicates that the corresponding row index in the target data file is dead.
  • Deletion vectors are extremely small, often consuming only a few kilobytes even when tracking thousands of deleted rows.
  • In Delta Lake, multiple deletion vectors for several data files can be co-located or packed into small shared storage files to prevent small file clutter.

How deletion vectors speed up operations

In traditional table operations, deleting a single record required rewriting a 200 MB Parquet file:

  • With deletion vectors enabled, a DELETE command finishes in seconds: the engine finds the matching row positions, writes a small deletion vector file, and commits a reference to that vector in the transaction log.
  • An UPDATE operation works similarly by writing a deletion vector for the old row position and appending a new record to a fresh Parquet file.
  • This pattern turns expensive file rewrites into lightweight append-only metadata operations.

Compaction and physical deletion

Because base data files still physically hold deleted records on disk, space is not reclaimed immediately. Over time, if a data file accumulates too many deletion vectors, reader decompression overhead increases. Data platform teams run maintenance commands like OPTIMIZE or periodic compaction jobs. These background jobs rewrite data files that exceed a threshold of deleted rows, physically stripping dead rows and removing obsolete deletion vectors.

PreviousNext