Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Orchestration & Reliable Pipelines

Progress0/14
x

What is a Pipeline

  • What is a data pipeline?12m
  • Scheduling and cron12m
  • Monitoring and alerting12m

DAG Mental Model

  • Why a scheduler12m
  • Tasks, edges, topological order14m
  • Operators, sensors, and XCom12m

Retries, Idempotency, Dedup

  • Retries in the DAG runner14m
  • Idempotent loads14m
  • Deduplicate before load12m

Backfills, Catchup, Intervals

  • Backfills and catchup12m
  • Data interval vs execution clock12m

CDC Concepts

  • CDC: inserts, updates, deletes12m
  • Apply CDC idempotently14m

Capstone

  • Capstone: a reliable night job16m
Back to track
  1. Learn
  2. Orchestration & Reliable Pipelines
  3. What is a Pipeline
  4. Monitoring and alerting

Lesson 3 of 14 · Theory first, then run it

Monitoring and alerting

orchestrationpythonbeginner12 min

Overview

A run is green or red. Logs explain why. An alert fires when the DAG is red or an SLA is late. This tab only counts failures.

On this page7 sections›
  1. 1Goal
  2. 2Why this order
  3. 3The checklist
  4. 4Worked pass
  5. 5Common beginner questions
  6. 6What comes next
  7. 7Practice

Goal

Monitoring is knowing whether last night's pipeline succeeded. Each run (and each task inside it) ends in a state: success or fail. The DAG is green when the required tasks succeeded. The DAG is red when one of them failed.

Logs are the text the job printed: the file it opened, the row count, the stack trace. When a DAG is red, you open the failed task and read the log. Alerts are messages (email, Slack, PagerDuty) that fire when the DAG is red, or when an SLA is missed: gold not ready by 08:00.

No Airflow UI lives in this tab. A list of run dicts with ok True or False is the model. You count the failures. On the job you will click the red square. The habit is the same: state first, log second, alert if nobody would notice otherwise.

Why this order

A pipeline that fails silently is worse than one that never ran. The dashboard still shows yesterday's GMV. Finance books the number. Two days later someone notices a gap. Monitoring exists so a red run is visible in minutes, not in a quarterly audit.

Alerts without logs create noise: a page that says 'DAG failed' and no link to the task. Logs without alerts create graves: the failure sat in a UI nobody opened. You want both, plus an SLA clock so 'it ran' is not confused with 'it was on time.'

The checklist

Think of a packing line with a status light. Green means the shift finished. Red means stop and read the clipboard (the log). An alert is the page that wakes someone if the light stays red past the SLA.

Run state, then log, then alert
DAG run startsEach task: green orredOpen the log on redAlert if SLA is atrisk

The scheduler records success or fail. A red task opens a log. An alert fires if nobody would see the red square in time.

Success/fail is a boolean on the task, not a vibe. A job that exited 0 with empty gold is still 'green' to the scheduler. Data quality gates (a later track) catch that lie. This lesson only counts runs whose ok flag is False.

State, logs, alerts, SLA. Four instruments. One night job.

SignalWhat it tells youWhat it does not tell you
Task state (green/red)Did the process finish without an error?Whether GMV is correct
LogsWhy it failed, what file, which exceptionWhether anyone saw the failure
AlertA human should look nowThe root cause (that is in the log)
SLAWas gold ready by the promised clock?Whether the DAG is currently running

An SLA is a promise: yesterday's gold.fct_purchases complete by 08:00 UTC. An SLI is the measurement: minutes of lag, or a boolean on_time. Alert on the miss. Do not page on every retry that later succeeded, or on-call will ignore you.

Worked pass

Two simulated runs: one succeeded, one failed. Count the failures. The Python editor is not a monitoring product. It is enough to practice 'how many are red.'

PythonFilter runs where ok is False
RUNS = [
    {"id": 1, "ok": True},
    {"id": 2, "ok": False},
]
failed = [row for row in RUNS if not row["ok"]]
print("red ids", [row["id"] for row in failed])
print("failed_count", len(failed))

A helper returns the integer the dashboard would show as 'failed last night.' Print it. Store it in result in the exercise.

PythonCount, do not only list
def failed_count(runs):
    return sum(1 for row in runs if not row["ok"])

print(failed_count(RUNS))

Common beginner questions

Is a green DAG enough?

No. Green means the process exited 0. Null order_ids can still land in gold. Quality gates and tests (Data Quality track, dbt tests) check the data. Monitoring checks the run. You need both.

Should every retry page me?

Usually no. Retry extract three times, then fail the task, then alert. Paging on attempt 1 trains people to ignore the channel. Alert on terminal failure and on SLA miss.

Where do logs live in production?

In the orchestrator UI (Airflow task log), in Cloud Logging / CloudWatch / Azure Monitor, and sometimes in a warehouse table of run metadata. Start with the red task's log. Do not grep a laptop.

Do not alert only on 'I happened to look'

If the only monitor is a human opening the UI at 09:00, weekend failures sit until Monday. Attach an alert to the DAG's failed state and to the SLA clock.

Green is not correct

The Quality track's first lesson is that a green DAG with null keys is not a successful night. This lesson only teaches you to see the red square. Bring both instincts to production.

What comes next

The next module is DAG Mental Model: why a scheduler exists beyond cron, how tasks depend on each other, and Airflow's vocabulary, still simulated in the Python editor with no Airflow process.

Practice

Run Sample to count failed runs. Then complete Exercise: with one success and one failure in RUNS, assign 1 to result and print it.

Practicals · load into the editor

After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.

Rate:
Was this useful?
Scheduling and cronWhy a scheduler