Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is data skipping?

File formats & storage · Performance & Layout

What is data skipping?

Mediumformat-13
data skippingstatisticsmin maxParquetperformance

Question

What is data skipping in columnar lakes and table formats?

Solution

Data skipping means the engine uses stored statistics (and sometimes indexes/bloom filters) to avoid reading files, row groups, or pages that cannot satisfy a query filter.

Building blocks

  • File / row-group min and max per column
  • Null counts, distinct counts (when present)
  • Bloom filters (ORC/Parquet/some table formats)
  • Partition values and table-format manifests
Query: WHERE event_date = '2024-01-02' AND amount > 1000

1) Partition pruning   → drop other date folders
2) Manifest / file prune → drop files whose stats cannot match
3) Row-group skip      → inside remaining Parquet files
4) Decode only needed columns (column pruning)

What makes skipping effective

  • Selective predicates on columns that have tight min/max ranges
  • Good layout: partitioning, clustering, Z-order, sorted writes
  • Avoiding huge files where every column spans the full domain

Interview tip: Data skipping is the umbrella term; partition pruning, predicate pushdown, and Z-order are ways to make skipping actually happen.

PreviousNext