Learn · orchestration
How pipelines run themselves, in order, and recover when a step fails.
Why pipelines need schedulers, how DAGs work, retries, idempotent loads, backfills, and CDC. Simulated runner (no Airflow scheduler in this tab).
Core foundational modules are free. Advanced production modules need Pro.
Playable walkthroughs: watch the system move, predict the next step, stamp a memory seal, then practice. Completing a Trace counts toward readiness.
You have three scripts. Script B needs A to finish, C needs B, and all three must run after the vendor file lands, which is usually 2 a.m. but sometimes 5 a.m. You start with cron and a sleep. Then a step fails halfway, and rerunning the whole chain double-loads yesterday's data. Then someone asks you to reprocess last March. Orchestration is the answer to all three problems: dependencies, retries that are safe, and reruns that produce the same answer.
Nobody hires you to write one script. They hire you to own the thing that runs at 2 a.m., and to be the person who knows what to do when it goes red. Most of that skill is dependency thinking, idempotency, and backfills.
DAGs are Python. You need functions and imports, nothing exotic.
Some pipeline experience
Having written a script that reads, transforms, and writes makes every lesson here concrete. The Core Python capstone is enough.
6 stages, in the order they build on each other. Each stage lists the modules and lessons it covers, and what you should be able to do by the end of it.
Start with the pain: cron plus sleep, and what breaks. Without feeling that, a scheduler looks like unnecessary machinery.
What is a PipelineFree
0/3
What a data pipeline is, why we schedule it, and how to know when it fails.
By the end of this stage
You can explain what a scheduler gives you that cron does not.
A pipeline is a graph of dependencies, not a list of steps. Once you see it that way, parallelism and failure isolation become obvious.
DAG Mental ModelFree
0/3
Why a scheduler exists, how tasks depend, and Airflow's vocabulary without an Airflow process.
Module Checkpoint
1 conceptual questions to verify mastery
By the end of this stage
You can take a tangle of scripts and draw the DAG, including what can safely run in parallel.
This is the core of the track. A retry is only safe if running the step twice leaves the data the same as running it once, and most naive loads fail that test.
Retries, Idempotency, Dedup
0/3
The parts you can actually write and test in this tab: retry a flaky call, load once, drop duplicate keys.
Module Checkpoint
2 conceptual questions to verify mastery
By the end of this stage
You can write a load that is safe to rerun, and explain the difference between at least once and exactly once delivery.
Where people get stuck
If you take one thing from this track, take idempotency. It is the most asked and least understood topic in pipeline interviews.
Pipelines are about time windows, and time is where the subtle bugs live: which interval a run belongs to, what happens when you reprocess history, whether late data lands in the right partition.
Backfills, Catchup, Intervals
0/2
Rerunning last week, catching up after a pause, and the difference between when it ran and which data it covers.
Module Checkpoint
1 conceptual questions to verify mastery
By the end of this stage
You can run a backfill over a date range without double counting or leaving gaps.
Copying a whole table every night stops working as it grows. Capturing only what changed is the standard answer, and it brings its own problems: deletes, ordering, and schema drift.
CDC Concepts
0/2
Change data capture is a stream of inserts/updates/deletes. Apply it idempotently, the same habit as SCD Type 2.
By the end of this stage
You can explain how CDC works and what it costs compared to a full reload.
Put it together: a dependency graph with retries, idempotent loads, and a backfill you can defend.
Capstone
0/1
Topo order, retries on extract, idempotent load, dedup. The night job in one script.
By the end of this stage
You have a pipeline design you can talk through end to end, including what happens when it fails.
Finishing the lessons is not the goal. These are the things you should be able to do afterwards, and each one is worth checking honestly.
30 minutes a day
about 6 sessions
1 hour a day
about 3 sessions
4 hours a weekend day
about 1 session
The scheduler in this track is simulated in your browser, so there is no Airflow to install and nothing to break. The concepts transfer directly to Airflow, Dagster, or Prefect. Focus on the reliability module rather than tool syntax, because the tools change and the reasoning does not.
These counts cover reading and the built-in exercises only. Real practice on the drills and a capstone will add to it, and that time is where most of the learning happens.
| Common mistake | What to do instead |
|---|---|
| Learning Airflow syntax before learning dependency and idempotency thinking. | The syntax takes a day. The reasoning is what makes your pipelines survive, and it is what gets tested. |
| Turning on retries without checking whether the step is safe to repeat. | Retries plus a non-idempotent insert equals duplicate rows. Make the write safe first. |
| Treating backfill as 'just run it again for old dates'. | Backfills interact with partitions, late data, and downstream jobs. Plan the range and the overwrite semantics. |