When a task instance remains stuck in the queued state, it means the Airflow scheduler has determined the task is ready to run and placed it on the executor's backlog, but the executor has not successfully assigned it to an active worker process.
Step 1: Check concurrency limits and pools
Start with the cheapest and most common cause by inspecting the Airflow web UI:
- Pool saturation: Check the Pools tab. If the task is assigned to a custom pool whose slots are 100% in use, the task waits in
queueduntil running tasks complete and release slots. - DAG concurrency: Check
max_active_taskson the DAG and the globalparallelismsetting in Airflow configuration. If the cluster has reached its global concurrency cap, new tasks stay queued.
Step 2: Verify queue routing and worker availability
If using CeleryExecutor:
Investigation path for CeleryExecutor: 1. Check task definition: Does it specify a custom queue? (e.g., queue="gpu_tasks") 2. Check Celery workers: Are any active workers subscribed to that specific queue? 3. Check message broker: Can the scheduler connect to Redis or RabbitMQ?
If a task specifies queue="heavy", but all Celery worker processes were started listening only to the default queue, the task message sits in the broker indefinitely. No worker will ever claim it.
Step 3: Inspect Kubernetes pod scheduling
If running on KubernetesExecutor:
Check the Kubernetes cluster using kubectl get pods -n airflow. Look for worker pods in the Pending state. Run kubectl describe pod <pod-name> to inspect Kubernetes events. Common causes include:
- Insufficient cluster CPU or memory to satisfy the task pod's resource requests.
- Node auto-scalers taking several minutes to provision new virtual machines.
ImagePullBackOffcaused by invalid container image tags or registry authentication failures.
Step 4: Verify scheduler loop health
Inspect scheduler logs and database connection latency. If the scheduler is experiencing metadata database lock contention, it can place tasks into the database queue table but fail to dispatch executor commands.
Prevention
To prevent queue stalls from impacting SLAs:
- Set up Prometheus alerts on
airflow_task_queued_duration_secondsto notify the team when any task sits queued longer than 15 minutes. - Enforce standard resource requests on Kubernetes pods to avoid unbounded resource allocation.
- Avoid orphaned Celery queue names by maintaining queue definitions in centralized deployment configs.