Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. max_active_runs, max_active_tasks and concurrency

Airflow & DAGs · Scheduling Deep Dive

max_active_runs, max_active_tasks and concurrency

Mediumairflow-46
concurrencypoolsschedulerresource-management

Question

What do max_active_runs, max_active_tasks and pools each limit?

Solution

Airflow provides concurrency controls across several levels of its hierarchy. Configuring them properly prevents pipeline runs from exhausting worker threads, overloading the scheduler, or bringing down production databases.

Concurrency settings by scope

Each setting restricts task parallelization within a specific boundary:

  • max_active_runs: Defined per DAG. Limits how many concurrent DAG run instances can be active simultaneously. When backfilling 60 days, setting max_active_runs=1 forces runs to execute sequentially rather than launching 60 runs at once.
  • max_active_tasks: Defined per DAG (formerly named concurrency). Caps how many total task instances can run in parallel across all active runs of that single DAG.
  • pools: Defined globally across the entire Airflow deployment. Tasks assign themselves to a pool with a fixed slot count.
  • parallelism: Global setting in airflow.cfg. Sets the absolute ceiling on concurrent running task instances across all DAGs in the cluster.

Protecting external databases with pools

While max_active_tasks limits concurrency inside one DAG, an external database is often queried by multiple independent DAGs. If ten pipelines run simultaneously, each opening five database connections, the source database can crash under connection limits.

# Tasks from different DAGs share the same pool allocation
extract_customers = PostgresOperator(
    task_id="extract_customers",
    pool="postgres_read_pool",
    sql="SELECT * FROM customers;",
)

By creating postgres_read_pool in the Airflow UI with 4 slots, Airflow limits concurrent database queries to 4 across all DAGs. Remaining tasks stay in the queued state until a slot frees up.

Worker concurrency

Celery and Kubernetes executors enforce their own infrastructure boundaries. In Celery, worker_concurrency determines how many worker processes a single Celery node runs. If a task is CPU-intensive, setting worker concurrency too high leads to CPU throttling and degraded execution speed.

PreviousNext