Partition pruning means Spark skips entire storage partitions that cannot match the query filters. If data is laid out as dt=.../region=..., a filter on those columns can avoid reading most files.
Example
Dataset layout: /events/dt=2024-01-01/... /events/dt=2024-01-02/... /events/dt=2024-01-03/... Query: WHERE dt = '2024-01-02' Spark reads only dt=2024-01-02 directories (partition pruning)
df = spark.read.parquet("/lake/events")
df.filter(df.dt == "2024-01-02").count()Requirements
- Partition columns must be present in the path/metastore metadata
- Filters must be selective and applied in a way Catalyst can push to the data source
- Dynamic partition pruning can also prune partitions of a fact table using values discovered from a dimension side at runtime (broadcast/DPP)
Interview tip
Pair with column pruning (read only needed Parquet columns) and predicate pushdown for nested filters inside files.