Introduced in Airflow 2.7, setup and teardown tasks provide a native syntax for managing the lifecycle of temporary infrastructure, guaranteeing that cleanup tasks execute even when intermediate pipeline tasks fail.
The fragility of legacy cleanup patterns
Before setup and teardown tasks, spinning up an ephemeral cluster (like an EMR or Dataproc cluster) and destroying it required configuring a cleanup task with trigger_rule="all_done". This pattern suffered from two major flaws:
First, if the cluster creation task failed at the very start, the cleanup task still ran under all_done, attempting to delete a cluster that was never created and throwing confusing secondary errors.
Second, if an intermediate transformation task failed, the cleanup task ran and succeeded, which often marked the overall DAG run as success unless complicated status tracking was implemented.
The setup and teardown syntax
Modern Airflow provides the .as_setup() and .as_teardown() methods:
from airflow.decorators import dag, task
@dag(schedule="@daily", start_date=datetime(2025, 1, 1))
def compute_pipeline():
@task
def create_cluster():
return "cluster_id_101"
@task
def run_spark_job(cluster_id):
# Heavy data processing
pass
@task
def delete_cluster(cluster_id):
# Tears down ephemeral infrastructure
pass
setup_task = create_cluster().as_setup()
work_task = run_spark_job(setup_task)
teardown_task = delete_cluster(setup_task).as_teardown(setups=setup_task)
setup_task >> work_task >> teardown_taskKey execution behaviors
Setup and teardown tasks enforce clear lifecycle rules:
- If
create_clusterfails during setup, both the work task and the teardown task are automatically skipped, avoiding pointless cleanup attempts on non-existent infrastructure. - If
run_spark_jobcrashes,delete_clusteris guaranteed to execute, preventing abandoned cloud compute resources from running up massive bills overnight. - The teardown task's completion does not mask failures: if
run_spark_jobfails, the overall DAG run is correctly recorded as failed.