Predicate pushdown means the engine pushes filter conditions (WHERE predicates) as close to the data source as possible, so less data is read, decoded, and shipped to compute.
Problem first. Without pushdown, Spark might open every Parquet row group, decode columns, then filter amount > 100 in memory. With pushdown, the Parquet reader uses min/max (and other stats) to skip whole chunks that cannot match.
Query: WHERE amount > 100 AND status = 'PAID'
Logical plan has Filter
|
Engine pushes predicates into Parquet/ORC/JDBC scan
|
Reader skips non-matching row groups / pages / partitionsWhere it shows up
- Parquet/ORC: row-group / stripe / page stats
- JDBC: generate SQL
WHEREin the database - Partition filters: skip directories (related idea: partition pruning)
What blocks pushdown
- Opaque Python UDFs wrapping columns before the filter
- Expressions the source cannot understand
- Row-ish text formats with weak stats (plain CSV)
Interview tip: Pair predicate pushdown with column pruning and partition pruning. Say: "push filters and projections to the source; keep columnar files so stats work."