Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Bloom filters in Parquet and table formats

File formats & storage · Storage Layout

Bloom filters in Parquet and table formats

Mediumfile-formats-42
bloom-filtersparquetindexingdata-skippingperformance

Question

What are Bloom filters, and when do they help in data lakes?

Solution

A Bloom filter is a space-efficient probabilistic data structure that tests whether an element is definitely not present in a dataset or might be present. In data lakes, Bloom filters accelerate point lookup queries on high-cardinality columns (such as uuid, customer_id, or device_mac_address) where standard min and max statistics fail to prune files because every data file spans almost the entire value range. Bloom filters are natively supported in Parquet file footers, Databricks Delta Lake, and Apache Iceberg, but they incur additional storage space and compute time during writes.

Probabilistic filtering mechanics

To understand why Bloom filters are valuable, consider how standard file pruning works:

  • Every Parquet column chunk records its minimum and maximum values. If a query filters WHERE order_date = '2025-01-01', and a file's range is February 2025, the reader skips the file immediately.
  • However, if a query filters for a single UUID like WHERE session_id = 'e4a2-91f8', min/max pruning is useless because string UUIDs are randomly distributed across every file. Every file has minimums starting with 00 and maximums starting with ff.
  • During file writes, column values are hashed across multiple hash functions to set bits in a small bit array.
  • When an engine executes a point query, it hashes the search key and checks the Bloom filter bit array for that file or row group.
  • If any checked bit is 0, the engine knows with 100 percent certainty that the key does not exist in that file, skipping the file entirely without reading data pages.
  • If all bits are 1, the key might be in the file, prompting the engine to read the file.

Point lookup acceleration on unique keys

Bloom filters shine in dimension lookups and point queries:

  • A query looking up a single transaction out of 10 billion rows can skip 99 percent of data files using Bloom filters even when the table is not partitioned by transaction ID.
  • This drastically reduces read I/O and query latency on high-cardinality primary key lookups.

Cost and maintenance overhead

However, Bloom filters should be applied selectively:

  • Generating Bloom filters requires extra CPU hashing time during write operations.
  • Bloom filter structures add 1 to 5 percent overhead to storage file sizes.
  • They provide zero value for range queries (such as amount > 500) and are redundant on low-cardinality columns where dictionary encoding already tracks distinct keys.
  • Reserve Bloom filters specifically for selective point-lookup ID columns in high-value query paths.
PreviousNext