Column pruning (also called projection pushdown) means the engine reads only the columns required by the query, not every column in the file.
Problem first. A wide fact table might have 80 columns. A dashboard needs order_id, amount, and region. In a row store or SELECT * CSV mindset, you still pay for the other 77 fields. In Parquet/ORC, those other column chunks can stay on disk.
File columns: id | user | amount | notes | json_blob | ... Query needs: amount, region Column pruning → open only those column chunks
# Good: project early so the scan can prune columns
df = spark.read.parquet("/lake/orders").select("order_id", "amount", "region")
# Bad habit: pull everything then drop later
df = spark.read.parquet("/lake/orders") # may still prune if plan is smart,
df = df.select("order_id") # but SELECT * patterns and UDFs hurtWhy it matters financially
Less I/O, less CPU decompression, less memory, cheaper cloud scans.
Interview tip: "Column pruning = do not read unused columns." Contrast with predicate pushdown (skip rows/chunks) and partition pruning (skip directories).