Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is partition pruning?

File formats & storage · Extra High-Value

What is partition pruning?

Mediumformat-18
partition pruningHive partitionsSparkfilters

Question

What is partition pruning?

Solution

Partition pruning means the query engine skips entire storage partitions that cannot match the filter. If data lives under dt=... directories (or equivalent table-format partitions), a filter on dt avoids listing and reading the rest.

/events/dt=2024-01-01/...
/events/dt=2024-01-02/...
/events/dt=2024-01-03/...

WHERE dt = '2024-01-02'
→ only open dt=2024-01-02 (partition pruning)
df = spark.read.parquet("/lake/events")
df.filter(df.dt == "2024-01-02").count()

Requirements

  • Partition columns must be known to the metastore / path / table metadata
  • Filters must be selective and recognizable (avoid wrapping partition cols in opaque UDFs)
  • Dynamic partition pruning can prune a fact table using values discovered from a dimension at runtime

How it differs from related ideas

  • Partition pruning: skip directories / partitions
  • Predicate pushdown: skip row groups/pages *inside* files
  • Column pruning: skip unused columns

Interview tip: Always give a dated path example, then say pruning only helps if you filter on the partition columns.

PreviousNext