Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is column pruning?

File formats & storage · Formats & Table Formats

What is column pruning?

Easyformat-09
column pruningprojectionParquetI/O

Question

What is column pruning (projection pushdown)?

Solution

Column pruning (also called projection pushdown) means the engine reads only the columns required by the query, not every column in the file.

Problem first. A wide fact table might have 80 columns. A dashboard needs order_id, amount, and region. In a row store or SELECT * CSV mindset, you still pay for the other 77 fields. In Parquet/ORC, those other column chunks can stay on disk.

File columns: id | user | amount | notes | json_blob | ...
Query needs:          amount, region

Column pruning → open only those column chunks
# Good: project early so the scan can prune columns
df = spark.read.parquet("/lake/orders").select("order_id", "amount", "region")

# Bad habit: pull everything then drop later
df = spark.read.parquet("/lake/orders")  # may still prune if plan is smart,
df = df.select("order_id")               # but SELECT * patterns and UDFs hurt

Why it matters financially

Less I/O, less CPU decompression, less memory, cheaper cloud scans.

Interview tip: "Column pruning = do not read unused columns." Contrast with predicate pushdown (skip rows/chunks) and partition pruning (skip directories).

PreviousNext