Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Pipeline failure handling

Data platform · Pipelines

Pipeline failure handling

Mediumplatform-06
failuresretriesalertingreliability

Question

How should a data pipeline behave when a step fails?

Solution

Failures are normal. A production-minded answer covers detection, isolation, retries, and recovery.

Principles

  • Fail loudly on bad data that would corrupt marts
  • Retry transient errors (network, throttling) with backoff
  • Do not retry forever on poison records; quarantine them
  • Keep upstream outputs durable so you can resume
  • Alert on SLA breaches (freshness / lag), not only hard crashes
extract OK -> transform FAIL
  -> do not mark DAG success
  -> page on-call if SLA risk
  -> fix / replay from last good checkpoint

Tiny example

Airflow task writing a partition fails mid-write. Next run should overwrite the incomplete partition, not append.

Interview tip: Talk about checkpoints, idempotent writes, dead-letter queues, and clear ownership of alerts.

PreviousNext