Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Orchestration & Reliable Pipelines

Progress0/14
x

What is a Pipeline

  • What is a data pipeline?12m
  • Scheduling and cron12m
  • Monitoring and alerting12m

DAG Mental Model

  • Why a scheduler12m
  • Tasks, edges, topological order14m
  • Operators, sensors, and XCom12m

Retries, Idempotency, Dedup

  • Retries in the DAG runner14m
  • Idempotent loads14m
  • Deduplicate before load12m

Backfills, Catchup, Intervals

  • Backfills and catchup12m
  • Data interval vs execution clock12m

CDC Concepts

  • CDC: inserts, updates, deletes12m
  • Apply CDC idempotently14m

Capstone

  • Capstone: a reliable night job16m
Back to track
  1. Learn
  2. Orchestration & Reliable Pipelines
  3. What is a Pipeline
  4. What is a data pipeline?

Lesson 1 of 14 · Theory first, then run it

What is a data pipeline?

orchestrationpythonbeginner12 min

Overview

Extract, transform, and load as automated steps. A script you run by hand is not a pipeline you can trust overnight.

On this page7 sections›
  1. 1The idea
  2. 2Why this exists
  3. 3Picture this
  4. 4A small example
  5. 5Common beginner questions
  6. 6What comes next
  7. 7Practice

The idea

A data pipeline is a series of automated steps that move data from where it is born to where it is used. The usual three names are extract (pull from a source), transform (clean and shape), and load (write to a lake or warehouse). Together they are often called ETL.

Extract might call an orders API or download a vendor CSV. Transform might parse dates, drop duplicates, and join a product catalog. Load might write Parquet to bronze or INSERT into a warehouse table. Each step is a function you could run by hand. A pipeline is those functions wired together so they run without a person at the keyboard.

No Airflow scheduler runs in this tab. The Python editor practices the three words as a list. Later lessons add a clock, retries, and a graph. Passing the exercise does not schedule anything on a real machine.

Why this exists

Hand-running a script works on a Tuesday afternoon when you remember. It fails when the API is slow at 2 AM, when you skip the load because the extract looked empty, or when two people run the same script and double-count GMV. Dashboards then show the wrong revenue, and nobody can say which run was the real one.

Companies depend on last night's numbers at 08:00. A pipeline is how those numbers arrive on time, in order, without someone living on an alarm clock. Data engineering job postings assume you can describe extract, transform, and load before they ask about Airflow.

Picture this

Think of a conveyor on a loading dock. Each station does one job: inbound (extract), inspect and relabel (transform), put away (load). A person carrying one box is not a conveyor. The conveyor is the same path every night.

Extract, transform, load
Source (API or file)ExtractTransformLoadLake or warehouse

Data moves left to right. Transform should not start until extract finished. Load should not start until transform finished.

Order matters. Loading before extract finishes writes yesterday's file or an empty file. Transforming before load is fine if transform writes a new file; transforming the warehouse table before the load lands is how gold goes stale. Later you will draw that order as a DAG. For now, remember the three names in sequence.

  1. Extract: pull bytes from a source (API, database, SFTP, object storage).
  2. Transform: clean types, filter junk, join reference data, apply business rules.
  3. Load: write the result to durable storage the next consumer can read.
  4. Repeat on a schedule or when a new file lands. Do not wait for a human to press Run.

Automation is not extra. It is how the three steps stay in order overnight.

StepTypical inputTypical outputWhat goes wrong by hand
ExtractVendor API or CSV dropRaw file in bronzeForgot to run; API timeout ignored
TransformBronze fileCleaned rowsRan on an empty extract; duplicated keys
LoadCleaned rowsWarehouse table or gold fileLoaded twice; GMV doubled

ELT is a cousin: extract, load the raw file, then transform inside the warehouse with SQL. The three verbs are the same. The place transform runs (Spark job versus warehouse SQL) is the difference. This lesson does not force one spelling. It forces you to name the steps.

Trace the order

Walk the same messy partner CSV through ETL, then ELT. Predict the next step before it is shown. The Python list in Practice still waits at the end.

A small example

Write the three steps as a list. The Python editor cannot call a vendor API. It can make the vocabulary something you print before you meet a scheduler.

PythonName ETL in order
STEPS = ["extract", "transform", "load"]
print("a pipeline is")
for name in STEPS:
    print("-", name)
print("count", len(STEPS))

A tiny simulation runs three functions in order and records what each returned. On a real orchestrator, each function becomes a task with its own retries. Here it is still one script so you can see the sequence.

PythonThree functions, one sequence
def extract():
    return {"orders": 3, "source": "api"}

def transform(raw):
    return {"orders": raw["orders"], "status": "clean"}

def load(clean):
    return {"written": clean["orders"], "target": "warehouse"}

raw = extract()
clean = transform(raw)
written = load(clean)
print(raw, clean, written)

Common beginner questions

Is a Jupyter notebook a pipeline?

A notebook is a good place to explore. A pipeline is code that runs unattended, with a clear start, a clear output, and a record of whether it worked. Promote the notebook to a script (or a DAG) before you trust it at 2 AM.

Do I always transform before load?

ETL does. ELT loads first, then transforms in SQL. Both are pipelines. What you must not do is load gold that nobody transformed, or transform a table that this run never loaded.

Why not cron a single script?

For one file, that can work. When extract must retry, load must wait, and you need last Tuesday's log, a scheduler (the rest of this track) earns its keep. This lesson is the payload that scheduler will run.

A script you must remember is not overnight-safe

If the runbook is 'open the laptop and press Run,' the first vacation is an incident. Automate the three steps before you add fancy tools.

Before you start

This track assumes you completed Core Python. If you have not, start there first. Core Python ingest is the extract-and-clean you will later hang on a schedule.

What comes next

The next lesson is scheduling: how cron expressions say '2 AM daily,' and why a clock is not the same as a pipeline that retries and waits on dependencies.

Practice

Run Sample to print extract, transform, load. Then complete Exercise: assign that list to result and print it.

Practicals · load into the editor

After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.

Practice this

Same ideas as interview drills. These challenges open in the studio with a dataset and tests already set up.

  • Handle Missing ValuesInterview-style drill: Show dropna, mean fill, and ffill strategies labeled by method.Studiobeginnerpandas12 min
Rate:
Was this useful?
Orchestration & Reliable PipelinesScheduling and cron