Root-cause analysis finds the underlying reason a quality incident happened, not only the symptom ("dashboard revenue is wrong"). Good RCA prevents repeat failures.
A practical DE workflow
1. Detect: alert on freshness, volume, test failure, or stakeholder report. 2. Contain: pause publish / mark asset unhealthy / communicate stale-but-safe status. 3. Scope blast radius: use lineage: which marts, dashboards, and jobs are affected? 4. Localize the first bad hop: source empty? schema drift? bad join? late DAG? wrong filter in yesterday's deploy? 5. Confirm with evidence: reconciliations, partition diffs, git/deploy timeline, raw vs staging vs mart samples. 6. Fix + backfill: correct logic or upstream, reprocess impacted partitions idempotently. 7. Prevent: add a test/monitor/contract that would have caught this earlier; write a short postmortem.
Symptom: revenue_dashboard -30% Lineage: dashboard <- fct_orders <- stg_orders <- raw.orders Finding: raw.orders partition missing for Asia region after producer deploy Fix: restore producer feed, backfill partition, add source freshness + row-count anomaly gate
Interview tip: Emphasize contain → lineage → first bad hop → evidence → backfill → preventive test. RCA is a process, not a guess.