Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is predicate pushdown?

File formats & storage · Formats & Table Formats

What is predicate pushdown?

Mediumformat-08
predicate pushdownParquetfiltersoptimization

Question

What is predicate pushdown in file formats and query engines?

Solution

Predicate pushdown means the engine pushes filter conditions (WHERE predicates) as close to the data source as possible, so less data is read, decoded, and shipped to compute.

Problem first. Without pushdown, Spark might open every Parquet row group, decode columns, then filter amount > 100 in memory. With pushdown, the Parquet reader uses min/max (and other stats) to skip whole chunks that cannot match.

Query: WHERE amount > 100 AND status = 'PAID'

Logical plan has Filter
        |
Engine pushes predicates into Parquet/ORC/JDBC scan
        |
Reader skips non-matching row groups / pages / partitions

Where it shows up

  • Parquet/ORC: row-group / stripe / page stats
  • JDBC: generate SQL WHERE in the database
  • Partition filters: skip directories (related idea: partition pruning)

What blocks pushdown

  • Opaque Python UDFs wrapping columns before the filter
  • Expressions the source cannot understand
  • Row-ish text formats with weak stats (plain CSV)

Interview tip: Pair predicate pushdown with column pruning and partition pruning. Say: "push filters and projections to the source; keep columnar files so stats work."

PreviousNext