Pipeline reliability means consumers can trust that data arrives on time, complete, and correct, and that failures are detected and recoverable.
Reliable pipeline checklist Freshness -> data lands before the SLA clock Completeness-> expected partitions / row counts present Correctness -> tests, reconciliations, anomaly checks Recoverability -> idempotent replay / backfill path Observability -> metrics, logs, lineage, clear owners
Concrete practices
- Idempotent, partitioned writes
- Data quality tests (nulls, uniqueness, referential checks)
- Alerts on lag / failed runs / empty partitions
- Runbooks for common failures
- Separate "data landed" from "data certified" if needed
Reliability is not only uptime
A green Airflow DAG that loads wrong currency rates is still unreliable for the business.
Interview tip: Structure the answer as freshness + completeness + correctness + recoverability. That framework sounds senior even as a fresher.