Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Callbacks and alerting

Airflow & DAGs · Operating Airflow in Production

Callbacks and alerting

Mediumairflow-57
alertingcallbacksslackmonitoringon-failure-callback

Question

How do you get alerted when an Airflow task fails?

Solution

Airflow provides callback hooks that execute custom Python functions when tasks or DAG runs transition between states, allowing teams to deliver actionable alerts to Slack, Microsoft Teams, email, or PagerDuty.

Task and DAG callbacks

Airflow exposes callbacks across two execution scopes:

  • Task-level callbacks: on_failure_callback, on_success_callback, and on_retry_callback. These attach directly to individual operators or via default_args.
  • DAG-level callbacks: on_failure_callback and on_success_callback on the DAG object. These fire only when the overall DAG run finishes in a failed or completed state.
from airflow.providers.slack.notifications.slack import send_slack_notification

default_args = {
    # Fires a structured notification when all retries are exhausted
    "on_failure_callback": send_slack_notification(
        slack_conn_id="slack_default",
        text="Task {{ ti.task_id }} failed in DAG {{ dag.dag_id }} for {{ ds }}",
        channel="#data-pipeline-alerts",
    ),
}

Constructing actionable alerts

An effective alert must give the on-call engineer enough context to assess severity immediately without digging through logs. High-signal alerts include:

  • The pipeline name and the specific failing task identifier.
  • The logical_date or data interval being processed.
  • The error type or a brief snippet of the exception message.
  • A direct hyperlink to the Airflow task log URL (ti.log_url).

Alerting hygiene in production

A common mistake in early-career pipeline design is configuring alert webhooks on on_retry_callback. If an ingestion task retries three times due to transient network blips and then succeeds, sending an alert on every retry floods Slack channels with noise, causing alert fatigue. Alert only when a task fails permanently after all retries are exhausted, or at the DAG level when the entire run fails.

In Airflow 3, legacy SLAs were removed because they were unreliable and prone to false alerts. Teams now rely on DAG-level failure callbacks, custom deadline alert daemons, or Prometheus metrics to monitor pipeline completion windows.

PreviousNext