Data observability is continuous monitoring of whether data systems are healthy: freshness, volume, distribution, schema, lineage, and downstream impact.
Think of it as APM, but for tables and pipelines.
checks over time: row_count sudden drop? null_rate on email spiked? schema changed unexpectedly? dashboard SLA breached?
Practical stack pieces
- Pipeline metadata (Airflow/dbt run results)
- Table anomaly monitors
- Lineage to find blast radius
- On-call runbooks
Interview tip: List the pillars (freshness, volume, schema, distribution, lineage), then give one incident story shape: detect → triage lineage → fix → backfill.