Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Which failures should be retried?

Airflow & DAGs · Operating Airflow in Production

Which failures should be retried?

Mediumairflow-56
retriesbackofftimeoutsreliability

Question

How do you decide retries, retry_delay and timeouts for a task?

Solution

Setting task retries and timeouts correctly requires distinguishing between transient infrastructure hiccups and deterministic code failures. Retrying the former recovers pipelines automatically; retrying the latter wastes compute and delays incident response.

Transient versus deterministic failures

Configure retries based on error categories:

  • Retry transient errors: Network timeouts, HTTP 429 rate limits, database lock deadlocks, and spot virtual machine preemptions are temporary. A task configured with retries=3 and a 5-minute delay usually succeeds on subsequent attempts once the external system stabilizes.
  • Fail fast on deterministic errors: A SQL syntax error, a schema mismatch, missing table permissions, or an unparseable JSON file will never succeed on a second try. Retrying these errors simply burns data warehouse compute credits, delays alerts to the on-call engineer, and prolongs downstream SLA breaches.
# Production task retry configuration
sync_api = PythonOperator(
    task_id="sync_api",
    python_callable=fetch_from_api,
    retries=4,
    retry_delay=timedelta(minutes=2),
    retry_exponential_backoff=True,
    max_retry_delay=timedelta(minutes=30),
    execution_timeout=timedelta(hours=1),
)

Exponential backoff and maximum delays

When querying external APIs or shared databases, fixed retry intervals can exacerbate problems. If 20 tasks fail due to an API rate limit and all retry simultaneously after 60 seconds, they trigger another rate limit.

Setting retry_exponential_backoff=True doubles the wait duration after each failure (for example 2, 4, 8, then 16 minutes), giving external servers room to recover. Always combine backoff with max_retry_delay so the delay does not grow to unreasonable multi-hour waits.

Guarding against hung tasks

Every production task must define an execution_timeout. If a remote socket hangs without closing or a database lock causes a query to stall indefinitely, a task lacking a timeout occupies a worker slot forever.

Finally, remember that automatic retries are only safe if the underlying task is strictly idempotent. If a task writes records without deduplication or staging tables, retrying midway through creates duplicate records in production.

PreviousNext