Data skipping means the engine uses stored statistics (and sometimes indexes/bloom filters) to avoid reading files, row groups, or pages that cannot satisfy a query filter.
Building blocks
- File / row-group min and max per column
- Null counts, distinct counts (when present)
- Bloom filters (ORC/Parquet/some table formats)
- Partition values and table-format manifests
Query: WHERE event_date = '2024-01-02' AND amount > 1000 1) Partition pruning → drop other date folders 2) Manifest / file prune → drop files whose stats cannot match 3) Row-group skip → inside remaining Parquet files 4) Decode only needed columns (column pruning)
What makes skipping effective
- Selective predicates on columns that have tight min/max ranges
- Good layout: partitioning, clustering, Z-order, sorted writes
- Avoiding huge files where every column spans the full domain
Interview tip: Data skipping is the umbrella term; partition pruning, predicate pushdown, and Z-order are ways to make skipping actually happen.