Walk through debugging discipline: measure → hypothesize → fix → verify.
STAR outline
Situation: Nightly Spark/SQL job ran 3 hours and often missed the 8 AM SLA.
Task: Reduce runtime and stabilize success rate.
Action (pick what fits your story):
- Checked logs for OOM, skew, or full scans
- Profiled stages / EXPLAIN plan
- Fixed partition pruning, broadcast small dimension, or reduced shuffle
- Added retries with backoff and clearer failure alerts
- Split heavy transform from load
Result: Runtime down (e.g. 3h → 45m) or failure rate down; SLA met for N days.
Example talking points
> The join to a large fact was shuffling on an unpartitioned key. I filtered early, selected only needed columns, and broadcast the small lookup. I also added a freshness check so we fail loudly before dashboards go stale.
Fresher-honest version
College or personal project: "CSV ingest took 20 minutes locally because I loaded everything then filtered. I pushed filters earlier and chunked reads; it finished in under 3 minutes."
Interview tip: Name one root cause and one primary fix. Don't list ten micro-optimizations.