Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is partition pruning?

PySpark · Partitioning & Pruning

What is partition pruning?

Mediumpyspark-37
partition-pruningiooptimization

Question

What is partition pruning in Spark?

Solution

Partition pruning means Spark skips entire storage partitions that cannot match the query filters. If data is laid out as dt=.../region=..., a filter on those columns can avoid reading most files.

Example

Dataset layout:
  /events/dt=2024-01-01/...
  /events/dt=2024-01-02/...
  /events/dt=2024-01-03/...

Query: WHERE dt = '2024-01-02'
Spark reads only dt=2024-01-02 directories (partition pruning)
df = spark.read.parquet("/lake/events")
df.filter(df.dt == "2024-01-02").count()

Requirements

  • Partition columns must be present in the path/metastore metadata
  • Filters must be selective and applied in a way Catalyst can push to the data source
  • Dynamic partition pruning can also prune partitions of a fact table using values discovered from a dimension side at runtime (broadcast/DPP)

Interview tip

Pair with column pruning (read only needed Parquet columns) and predicate pushdown for nested filters inside files.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext