Failures are normal. A production-minded answer covers detection, isolation, retries, and recovery.
Principles
- Fail loudly on bad data that would corrupt marts
- Retry transient errors (network, throttling) with backoff
- Do not retry forever on poison records; quarantine them
- Keep upstream outputs durable so you can resume
- Alert on SLA breaches (freshness / lag), not only hard crashes
extract OK -> transform FAIL -> do not mark DAG success -> page on-call if SLA risk -> fix / replay from last good checkpoint
Tiny example
Airflow task writing a partition fails mid-write. Next run should overwrite the incomplete partition, not append.
Interview tip: Talk about checkpoints, idempotent writes, dead-letter queues, and clear ownership of alerts.