Partition pruning means the query engine skips entire storage partitions that cannot match the filter. If data lives under dt=... directories (or equivalent table-format partitions), a filter on dt avoids listing and reading the rest.
/events/dt=2024-01-01/... /events/dt=2024-01-02/... /events/dt=2024-01-03/... WHERE dt = '2024-01-02' → only open dt=2024-01-02 (partition pruning)
df = spark.read.parquet("/lake/events")
df.filter(df.dt == "2024-01-02").count()Requirements
- Partition columns must be known to the metastore / path / table metadata
- Filters must be selective and recognizable (avoid wrapping partition cols in opaque UDFs)
- Dynamic partition pruning can prune a fact table using values discovered from a dimension at runtime
How it differs from related ideas
- Partition pruning: skip directories / partitions
- Predicate pushdown: skip row groups/pages *inside* files
- Column pruning: skip unused columns
Interview tip: Always give a dated path example, then say pruning only helps if you filter on the partition columns.