Anomaly detection finds unusual changes in data metrics without hard-coding every bad case. Instead of only "amount must be >= 0," you watch whether today's volume, null rate, or sum is statistically weird vs history.
Why you need it
Static tests miss silent failures:
- Row count is non-zero but 70% lower than typical Mondays
emailnull rate jumps from 2% to 40%- Average order value doubles overnight after a currency bug
metric: daily_order_count history: Mon~10k, Tue~9k, ... today: 1.2k --> anomaly alert (even if "count > 0" still passes)
Common approaches
- Thresholds from rolling baselines (median ± k * MAD / stddev)
- Seasonality-aware models (weekday vs weekend)
- Simple rules first: WoW / DoD percent change
- ML detectors in commercial observability tools
Gotchas
- Seasonality and campaigns create false positives
- Cold start (new tables) has weak baselines
- Anomaly ≠ root cause; it is a signal to investigate
Interview tip: Contrast rule-based tests with statistical anomaly monitors. Give a volume or null-rate spike example, then say you still need lineage for RCA.