Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Orchestration & Reliable Pipelines

Progress0/14
x

What is a Pipeline

  • What is a data pipeline?12m
  • Scheduling and cron12m
  • Monitoring and alerting12m

DAG Mental Model

  • Why a scheduler12m
  • Tasks, edges, topological order14m
  • Operators, sensors, and XCom12m

Retries, Idempotency, Dedup

  • Retries in the DAG runner14m
  • Idempotent loads14m
  • Deduplicate before load12m

Backfills, Catchup, Intervals

  • Backfills and catchup12m
  • Data interval vs execution clock12m

CDC Concepts

  • CDC: inserts, updates, deletes12m
  • Apply CDC idempotently14m

Capstone

  • Capstone: a reliable night job16m
Back to track
  1. Learn
  2. Orchestration & Reliable Pipelines
  3. DAG Mental Model
  4. Why a scheduler

Lesson 4 of 14 · Theory first, then run it

Why a scheduler

orchestrationpythonbeginner12 min

Overview

A script you run by hand is not a pipeline. A scheduler owns the clock, retries, and run history.

On this page7 sections›
  1. 1The idea
  2. 2Why this exists
  3. 3Picture this
  4. 4A small example
  5. 5Common beginner questions
  6. 6What comes next
  7. 7Practice

The idea

A data pipeline is a series of steps that move and transform data: extract it from a source, load it into a warehouse, and transform it into something useful. When you have one script, you can run it by hand. When you have ten scripts that depend on each other and must run every night at 2 AM, you need something to manage the schedule, the order, and the failures. That something is an orchestrator.

An orchestrator (also called a scheduler or workflow engine) is software that runs your data jobs on a schedule, in the right order, with automatic retries when things go wrong. Popular orchestrators include Apache Airflow, Prefect, Dagster, and cloud-managed services like Google Cloud Composer or AWS MWAA.

This lesson is a simulation. There is no Airflow scheduler or cloud environment running in this tab. You practice the core ideas (tasks, schedules, retries, dependencies) in Python because those concepts transfer to every orchestrator you will use on the job.

Why this exists

Imagine a nightly data pipeline: a script pulls orders from an API, another script loads them into a database, and a third script builds a revenue report. If the API is slow and times out, but nothing enforces the order of steps, the load script can run anyway on an empty file, and the report shows zero revenue. All because nobody told the load step to wait until the extract step succeeded, and nobody told the extract step to try again.

Without an orchestrator, you rely on hope: hope that scripts run in the right order, hope that failures get noticed, hope that someone remembers what ran last Tuesday. An orchestrator replaces hope with guarantees.

Picture this

An orchestrator handles four jobs that a plain script cannot do well on its own.

The four jobs of a scheduler
Schedule (clock)Retry (resilience)Dependency (order)History (audit)

A plain script handles none of these reliably. An orchestrator handles all four.

  1. Schedule: start the pipeline at 2 AM every day without a human pressing a button.
  2. Retry: when an API call fails with a 503 error, wait 30 seconds and try again up to 3 times.
  3. Dependency: refuse to start the load step until the extract step has finished successfully.
  4. History: keep a record of every run, so you can answer 'what happened last Tuesday?' in seconds.

A cron job can handle the schedule. But cron cannot retry a failed task, enforce dependencies between tasks, or keep a browsable history. That is why orchestrators exist.

NeedPlain scriptOrchestrator
Start at 02:00Human alarm or cronBuilt-in timetable
API returns 503Script crashes silentlyRetries with exponential backoff
Load must wait for extractsleep(3600) and hopeDeclared dependency edge
What ran last Tuesday?grep through log filesBrowsable run history

A small example

Run the example below in this tab. Read the input, follow the code, then check the output matches what you expect.

Python
# What an orchestrator does, in simple Python
tasks = {
    "extract": {"depends_on": [], "retries": 3},
    "load":    {"depends_on": ["extract"], "retries": 2},
    "transform": {"depends_on": ["load"], "retries": 0},
}

def can_run(task_id, completed):
    """A task can run only when all its dependencies are done."""
    deps = tasks[task_id]["depends_on"]
    return all(d in completed for d in deps)

completed = set()
for task_id in ["extract", "load", "transform"]:
    if can_run(task_id, completed):
        print(f"Running {task_id}")
        completed.add(task_id)

The code above is a tiny model of what Airflow does internally. Each task declares its dependencies and retry count. The scheduler walks the list in order, only running a task when its upstream dependencies are satisfied.

Common beginner questions

Why can't I just use cron?

Cron handles the clock, but nothing else. It cannot retry a failed job, wait for a prior step to finish, or show you a dashboard of past runs. For a single independent script, cron is fine. For a multi-step pipeline, you need a real orchestrator.

Do I need Airflow specifically?

No. Airflow is the most widely used open-source orchestrator, so the vocabulary from this track (DAGs, tasks, operators) maps directly to it. But the concepts apply equally to Prefect, Dagster, or any managed service.

Is this track going to install Airflow?

No. You will practice the mental model and the Python patterns that every orchestrator shares. Installing Airflow on your own machine is a step you take outside this platform once you have the concepts.

Core Python retries apply here

You already practiced exponential backoff and retry logic in Core Python. A task inside an orchestrator is just a function with that same retry logic, placed on a graph.

sleep() is not a dependency

Using time.sleep(3600) to wait for a prior job is fragile. If the prior job takes longer than expected, your downstream job starts too early. Declared dependencies are the solution.

What comes next

In the next lesson, you will learn what a DAG is: the graph structure that orchestrators use to represent task dependencies. You will implement topological sort, the algorithm that determines the safe execution order.

Practice

Name the four jobs an orchestrator provides: schedule, retry, dependency, and history.

Practicals · load into the editor

After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.

Rate:
Was this useful?
Monitoring and alertingTasks, edges, topological order