Airflow provides concurrency controls across several levels of its hierarchy. Configuring them properly prevents pipeline runs from exhausting worker threads, overloading the scheduler, or bringing down production databases.
Concurrency settings by scope
Each setting restricts task parallelization within a specific boundary:
max_active_runs: Defined per DAG. Limits how many concurrent DAG run instances can be active simultaneously. When backfilling 60 days, settingmax_active_runs=1forces runs to execute sequentially rather than launching 60 runs at once.max_active_tasks: Defined per DAG (formerly namedconcurrency). Caps how many total task instances can run in parallel across all active runs of that single DAG.pools: Defined globally across the entire Airflow deployment. Tasks assign themselves to a pool with a fixed slot count.parallelism: Global setting inairflow.cfg. Sets the absolute ceiling on concurrent running task instances across all DAGs in the cluster.
Protecting external databases with pools
While max_active_tasks limits concurrency inside one DAG, an external database is often queried by multiple independent DAGs. If ten pipelines run simultaneously, each opening five database connections, the source database can crash under connection limits.
# Tasks from different DAGs share the same pool allocation
extract_customers = PostgresOperator(
task_id="extract_customers",
pool="postgres_read_pool",
sql="SELECT * FROM customers;",
)By creating postgres_read_pool in the Airflow UI with 4 slots, Airflow limits concurrent database queries to 4 across all DAGs. Remaining tasks stay in the queued state until a slot frees up.
Worker concurrency
Celery and Kubernetes executors enforce their own infrastructure boundaries. In Celery, worker_concurrency determines how many worker processes a single Celery node runs. If a task is CPU-intensive, setting worker concurrency too high leads to CPU throttling and degraded execution speed.