Stay calm, stop the bleed, understand what committed, then repair with an idempotent replay.
extract OK → transform FAIL → load not started
or
load started writing day=2026-09-05 → crash mid-writeStep-by-step response
1. Check orchestrator state: which task failed, error, attempt count. 2. Assess partial writes: is the target partition incomplete or uncommitted? 3. Do not mark success or silently continue downstream. 4. Fix root cause (bad file, OOM, permissions, poison row). 5. Replay safely: overwrite the partition / merge from last checkpoint. 6. Validate: row counts, null checks, compare to source. 7. Communicate: freshness impact to BI owners if SLA is at risk.
Design that makes this easy
Stage-then-publish, partitioned overwrites, durable raw zone, clear runbooks.
Interview tip: Interviewers want process + idempotent recovery, not "I would rerun the DAG" alone. Say what you check before hitting retry.