Monitoring answers: Did it run? Is data fresh? Is it correct enough? Is the system healthy?
Layers of monitoring 1) Orchestration: task success/fail, duration, SLA miss 2) Data: row counts, null rates, freshness, anomalies 3) Platform: Kafka lag, Spark stage time, warehouse slots/cost 4) Business: revenue reconciliation vs payments source
Useful signals
- Job failure / retry spikes
- Data arrival time vs SLA
- Empty or missing partitions
- Volume spikes/drops vs yesterday/week-ago baseline
- Consumer lag (streaming)
- DQ test failures
Alerting hygiene
Page humans for SLA risk and data corruption risk. Ticket noisy warnings. Include runbook links.
Tools (examples)
Airflow/Dagster UI, metrics (Prometheus/CloudWatch), DQ (dbt tests, Great Expectations), lineage (OpenLineage), log search.
Interview tip: Structure as pipeline health + data health + platform health. Mention freshness and row-count anomaly checks explicitly.