Overview
A run is green or red. Logs explain why. An alert fires when the DAG is red or an SLA is late. This tab only counts failures.
On this page7 sections
Goal
Monitoring is knowing whether last night's pipeline succeeded. Each run (and each task inside it) ends in a state: success or fail. The DAG is green when the required tasks succeeded. The DAG is red when one of them failed.
Logs are the text the job printed: the file it opened, the row count, the stack trace. When a DAG is red, you open the failed task and read the log. Alerts are messages (email, Slack, PagerDuty) that fire when the DAG is red, or when an SLA is missed: gold not ready by 08:00.
No Airflow UI lives in this tab. A list of run dicts with ok True or False is the model. You count the failures. On the job you will click the red square. The habit is the same: state first, log second, alert if nobody would notice otherwise.
Why this order
A pipeline that fails silently is worse than one that never ran. The dashboard still shows yesterday's GMV. Finance books the number. Two days later someone notices a gap. Monitoring exists so a red run is visible in minutes, not in a quarterly audit.
Alerts without logs create noise: a page that says 'DAG failed' and no link to the task. Logs without alerts create graves: the failure sat in a UI nobody opened. You want both, plus an SLA clock so 'it ran' is not confused with 'it was on time.'
The checklist
Think of a packing line with a status light. Green means the shift finished. Red means stop and read the clipboard (the log). An alert is the page that wakes someone if the light stays red past the SLA.
The scheduler records success or fail. A red task opens a log. An alert fires if nobody would see the red square in time.
Success/fail is a boolean on the task, not a vibe. A job that exited 0 with empty gold is still 'green' to the scheduler. Data quality gates (a later track) catch that lie. This lesson only counts runs whose ok flag is False.
State, logs, alerts, SLA. Four instruments. One night job.
| Signal | What it tells you | What it does not tell you |
|---|---|---|
| Task state (green/red) | Did the process finish without an error? | Whether GMV is correct |
| Logs | Why it failed, what file, which exception | Whether anyone saw the failure |
| Alert | A human should look now | The root cause (that is in the log) |
| SLA | Was gold ready by the promised clock? | Whether the DAG is currently running |
An SLA is a promise: yesterday's gold.fct_purchases complete by 08:00 UTC. An SLI is the measurement: minutes of lag, or a boolean on_time. Alert on the miss. Do not page on every retry that later succeeded, or on-call will ignore you.
Worked pass
Two simulated runs: one succeeded, one failed. Count the failures. The Python editor is not a monitoring product. It is enough to practice 'how many are red.'
RUNS = [
{"id": 1, "ok": True},
{"id": 2, "ok": False},
]
failed = [row for row in RUNS if not row["ok"]]
print("red ids", [row["id"] for row in failed])
print("failed_count", len(failed))A helper returns the integer the dashboard would show as 'failed last night.' Print it. Store it in result in the exercise.
def failed_count(runs):
return sum(1 for row in runs if not row["ok"])
print(failed_count(RUNS))Common beginner questions
Is a green DAG enough?
No. Green means the process exited 0. Null order_ids can still land in gold. Quality gates and tests (Data Quality track, dbt tests) check the data. Monitoring checks the run. You need both.
Should every retry page me?
Usually no. Retry extract three times, then fail the task, then alert. Paging on attempt 1 trains people to ignore the channel. Alert on terminal failure and on SLA miss.
Where do logs live in production?
In the orchestrator UI (Airflow task log), in Cloud Logging / CloudWatch / Azure Monitor, and sometimes in a warehouse table of run metadata. Start with the red task's log. Do not grep a laptop.
Do not alert only on 'I happened to look'
If the only monitor is a human opening the UI at 09:00, weekend failures sit until Monday. Attach an alert to the DAG's failed state and to the SLA clock.
Green is not correct
The Quality track's first lesson is that a green DAG with null keys is not a successful night. This lesson only teaches you to see the red square. Bring both instincts to production.
What comes next
The next module is DAG Mental Model: why a scheduler exists beyond cron, how tasks depend on each other, and Airflow's vocabulary, still simulated in the Python editor with no Airflow process.
Practice
Run Sample to count failed runs. Then complete Exercise: with one success and one failure in RUNS, assign 1 to result and print it.
Practicals · load into the editor
After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.