Tie optimization to a measurable before/after. Prefer one lever.
Common levers (DE interviews)
- Partition pruning / avoid full scans
- Column pruning (select only needed fields)
- File format (Parquet vs CSV)
- Cluster sizing / autoscaling down
- Incremental loads vs full reloads
- Caching only when reused
- Lifecycle policies on raw storage
STAR outline
Situation: Dev Spark cluster ran a wide CSV scan hourly; cloud bill spiked.
Action: Converted landing to Parquet partitioned by date, filtered to last 3 days, reduced executor count after measuring. Added cost dashboard note.
Result: Job runtime −60%, estimated compute cost −40%.
Fresher-honest version
"My laptop job OOMed on a big CSV. I switched to chunked Pandas reads and only kept needed columns; memory stayed flat."
Interview tip: Never optimize blindly. Say you measured first.