Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Anomaly detection for data

Data quality · Quality Foundations

Anomaly detection for data

Mediumquality-10
anomaly detectionbaselinesvolumenull rate

Question

What is anomaly detection in a data quality context?

Solution

Anomaly detection finds unusual changes in data metrics without hard-coding every bad case. Instead of only "amount must be >= 0," you watch whether today's volume, null rate, or sum is statistically weird vs history.

Why you need it

Static tests miss silent failures:

  • Row count is non-zero but 70% lower than typical Mondays
  • email null rate jumps from 2% to 40%
  • Average order value doubles overnight after a currency bug
metric: daily_order_count
history: Mon~10k, Tue~9k, ...
today:  1.2k  --> anomaly alert (even if "count > 0" still passes)

Common approaches

  • Thresholds from rolling baselines (median ± k * MAD / stddev)
  • Seasonality-aware models (weekday vs weekend)
  • Simple rules first: WoW / DoD percent change
  • ML detectors in commercial observability tools

Gotchas

  • Seasonality and campaigns create false positives
  • Cold start (new tables) has weak baselines
  • Anomaly ≠ root cause; it is a signal to investigate

Interview tip: Contrast rule-based tests with statistical anomaly monitors. Give a volume or null-rate spike example, then say you still need lineage for RCA.

PreviousNext