A runbook tells the person on call, who may not have built the pipeline, what the pipeline does and what to do when it breaks. It should be short enough to use at 3 a.m. and kept up to date.
What to include
- Purpose and owners: what the pipeline produces, who uses it, which team owns it, and how to reach them.
- Schedule and SLA: when it runs, how long it normally takes, and when the output must be ready.
- Inputs and outputs: upstream sources and downstream tables, dashboards and teams that depend on it, with links to lineage.
- Where things are: the repository, the orchestrator page, the logs, the dashboards, and the main tables.
- Common failures and fixes: a short list, each with the symptom, the likely cause, and the steps. Examples: "source file missing: check the partner SFTP, then contact X", "warehouse out of credits: ...", "schema change failed the check: ...".
- How to rerun and backfill safely: the exact commands or buttons, what parameters to set, and what must not be done (for example "do not run two backfills at once"). State whether the job is idempotent.
- How to know it is fixed: the checks to run, such as row counts or a freshness query.
- Escalation: who to call if the on-call person cannot solve it within a set time, and who to tell about the impact (business owners).
- Known quirks: planned maintenance windows of the source, or tables that are slow on month end.
Keep it useful
Store it next to the code (in the repository, as markdown), so changes to the pipeline and the runbook are reviewed together, and link it from every alert. Update it after each incident: if you had to find something out the hard way, write it down. Test it occasionally by asking a teammate who does not know the pipeline to follow it.
Format
Use headings and short steps, with real commands that can be copied. Long essays do not get read during an incident. A one-page layout is a good target.