Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Task stuck in queued

Airflow & DAGs · Operating Airflow in Production

Task stuck in queued

Hardairflow-53
scenariodebuggingqueued-statepoolstroubleshooting

Question

A task stays in queued state for a long time. What do you check?

Solution

When a task instance remains stuck in the queued state, it means the Airflow scheduler has determined the task is ready to run and placed it on the executor's backlog, but the executor has not successfully assigned it to an active worker process.

Step 1: Check concurrency limits and pools

Start with the cheapest and most common cause by inspecting the Airflow web UI:

  • Pool saturation: Check the Pools tab. If the task is assigned to a custom pool whose slots are 100% in use, the task waits in queued until running tasks complete and release slots.
  • DAG concurrency: Check max_active_tasks on the DAG and the global parallelism setting in Airflow configuration. If the cluster has reached its global concurrency cap, new tasks stay queued.

Step 2: Verify queue routing and worker availability

If using CeleryExecutor:

Investigation path for CeleryExecutor:
1. Check task definition: Does it specify a custom queue? (e.g., queue="gpu_tasks")
2. Check Celery workers: Are any active workers subscribed to that specific queue?
3. Check message broker: Can the scheduler connect to Redis or RabbitMQ?

If a task specifies queue="heavy", but all Celery worker processes were started listening only to the default queue, the task message sits in the broker indefinitely. No worker will ever claim it.

Step 3: Inspect Kubernetes pod scheduling

If running on KubernetesExecutor:

Check the Kubernetes cluster using kubectl get pods -n airflow. Look for worker pods in the Pending state. Run kubectl describe pod <pod-name> to inspect Kubernetes events. Common causes include:

  • Insufficient cluster CPU or memory to satisfy the task pod's resource requests.
  • Node auto-scalers taking several minutes to provision new virtual machines.
  • ImagePullBackOff caused by invalid container image tags or registry authentication failures.

Step 4: Verify scheduler loop health

Inspect scheduler logs and database connection latency. If the scheduler is experiencing metadata database lock contention, it can place tasks into the database queue table but fail to dispatch executor commands.

Prevention

To prevent queue stalls from impacting SLAs:

  • Set up Prometheus alerts on airflow_task_queued_duration_seconds to notify the team when any task sits queued longer than 15 minutes.
  • Enforce standard resource requests on Kubernetes pods to avoid unbounded resource allocation.
  • Avoid orphaned Celery queue names by maintaining queue definitions in centralized deployment configs.
PreviousNext