Use a systematic debug ladder. Interviewers want process under pressure.
Debug steps
1. Confirm the symptom: which metric, grain, time range, consumer 2. Check freshness: did the job run? late? failed? 3. Isolate layer: source vs landing vs transform vs BI 4. Compare: today's vs yesterday; warehouse vs source sample 5. Reproduce on a small partition 6. Fix forward + backfill if needed 7. Prevent: test/alert/runbook
Example outline (STAR-ready)
Situation: Marketing said new users dropped 40% overnight.
Action: Verified dashboard filter; checked Airflow success; compared warehouse count to auth DB; found a timezone cut-off bug in the incremental watermark. Corrected watermark, backfilled two days, added a day-over-day anomaly alert.
Result: Numbers restored; false alarm avoided next time.
Fresher tip
Practice this on a broken personal DAG: "read logs → check input file → validate transform → check output table."
Interview tip: Say what you would communicate during the outage (status + ETA).