Setting task retries and timeouts correctly requires distinguishing between transient infrastructure hiccups and deterministic code failures. Retrying the former recovers pipelines automatically; retrying the latter wastes compute and delays incident response.
Transient versus deterministic failures
Configure retries based on error categories:
- Retry transient errors: Network timeouts, HTTP 429 rate limits, database lock deadlocks, and spot virtual machine preemptions are temporary. A task configured with
retries=3and a 5-minute delay usually succeeds on subsequent attempts once the external system stabilizes. - Fail fast on deterministic errors: A SQL syntax error, a schema mismatch, missing table permissions, or an unparseable JSON file will never succeed on a second try. Retrying these errors simply burns data warehouse compute credits, delays alerts to the on-call engineer, and prolongs downstream SLA breaches.
# Production task retry configuration
sync_api = PythonOperator(
task_id="sync_api",
python_callable=fetch_from_api,
retries=4,
retry_delay=timedelta(minutes=2),
retry_exponential_backoff=True,
max_retry_delay=timedelta(minutes=30),
execution_timeout=timedelta(hours=1),
)Exponential backoff and maximum delays
When querying external APIs or shared databases, fixed retry intervals can exacerbate problems. If 20 tasks fail due to an API rate limit and all retry simultaneously after 60 seconds, they trigger another rate limit.
Setting retry_exponential_backoff=True doubles the wait duration after each failure (for example 2, 4, 8, then 16 minutes), giving external servers room to recover. Always combine backoff with max_retry_delay so the delay does not grow to unreasonable multi-hour waits.
Guarding against hung tasks
Every production task must define an execution_timeout. If a remote socket hangs without closing or a database lock causes a query to stall indefinitely, a task lacking a timeout occupies a worker slot forever.
Finally, remember that automatic retries are only safe if the underlying task is strictly idempotent. If a task writes records without deduplication or staging tables, retrying midway through creates duplicate records in production.