Overview
A script you run by hand is not a pipeline. A scheduler owns the clock, retries, and run history.
On this page7 sections
The idea
A data pipeline is a series of steps that move and transform data: extract it from a source, load it into a warehouse, and transform it into something useful. When you have one script, you can run it by hand. When you have ten scripts that depend on each other and must run every night at 2 AM, you need something to manage the schedule, the order, and the failures. That something is an orchestrator.
An orchestrator (also called a scheduler or workflow engine) is software that runs your data jobs on a schedule, in the right order, with automatic retries when things go wrong. Popular orchestrators include Apache Airflow, Prefect, Dagster, and cloud-managed services like Google Cloud Composer or AWS MWAA.
This lesson is a simulation. There is no Airflow scheduler or cloud environment running in this tab. You practice the core ideas (tasks, schedules, retries, dependencies) in Python because those concepts transfer to every orchestrator you will use on the job.
Why this exists
Imagine a nightly data pipeline: a script pulls orders from an API, another script loads them into a database, and a third script builds a revenue report. If the API is slow and times out, but nothing enforces the order of steps, the load script can run anyway on an empty file, and the report shows zero revenue. All because nobody told the load step to wait until the extract step succeeded, and nobody told the extract step to try again.
Without an orchestrator, you rely on hope: hope that scripts run in the right order, hope that failures get noticed, hope that someone remembers what ran last Tuesday. An orchestrator replaces hope with guarantees.
Picture this
An orchestrator handles four jobs that a plain script cannot do well on its own.
A plain script handles none of these reliably. An orchestrator handles all four.
- Schedule: start the pipeline at 2 AM every day without a human pressing a button.
- Retry: when an API call fails with a 503 error, wait 30 seconds and try again up to 3 times.
- Dependency: refuse to start the load step until the extract step has finished successfully.
- History: keep a record of every run, so you can answer 'what happened last Tuesday?' in seconds.
A cron job can handle the schedule. But cron cannot retry a failed task, enforce dependencies between tasks, or keep a browsable history. That is why orchestrators exist.
| Need | Plain script | Orchestrator |
|---|---|---|
| Start at 02:00 | Human alarm or cron | Built-in timetable |
| API returns 503 | Script crashes silently | Retries with exponential backoff |
| Load must wait for extract | sleep(3600) and hope | Declared dependency edge |
| What ran last Tuesday? | grep through log files | Browsable run history |
A small example
Run the example below in this tab. Read the input, follow the code, then check the output matches what you expect.
# What an orchestrator does, in simple Python
tasks = {
"extract": {"depends_on": [], "retries": 3},
"load": {"depends_on": ["extract"], "retries": 2},
"transform": {"depends_on": ["load"], "retries": 0},
}
def can_run(task_id, completed):
"""A task can run only when all its dependencies are done."""
deps = tasks[task_id]["depends_on"]
return all(d in completed for d in deps)
completed = set()
for task_id in ["extract", "load", "transform"]:
if can_run(task_id, completed):
print(f"Running {task_id}")
completed.add(task_id)The code above is a tiny model of what Airflow does internally. Each task declares its dependencies and retry count. The scheduler walks the list in order, only running a task when its upstream dependencies are satisfied.
Common beginner questions
Why can't I just use cron?
Cron handles the clock, but nothing else. It cannot retry a failed job, wait for a prior step to finish, or show you a dashboard of past runs. For a single independent script, cron is fine. For a multi-step pipeline, you need a real orchestrator.
Do I need Airflow specifically?
No. Airflow is the most widely used open-source orchestrator, so the vocabulary from this track (DAGs, tasks, operators) maps directly to it. But the concepts apply equally to Prefect, Dagster, or any managed service.
Is this track going to install Airflow?
No. You will practice the mental model and the Python patterns that every orchestrator shares. Installing Airflow on your own machine is a step you take outside this platform once you have the concepts.
Core Python retries apply here
You already practiced exponential backoff and retry logic in Core Python. A task inside an orchestrator is just a function with that same retry logic, placed on a graph.
sleep() is not a dependency
Using time.sleep(3600) to wait for a prior job is fragile. If the prior job takes longer than expected, your downstream job starts too early. Declared dependencies are the solution.
What comes next
In the next lesson, you will learn what a DAG is: the graph structure that orchestrators use to represent task dependencies. You will implement topological sort, the algorithm that determines the safe execution order.
Practice
Name the four jobs an orchestrator provides: schedule, retry, dependency, and history.
Practicals · load into the editor
After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.