Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Data downtime and incident metrics

Data quality · Incidents & Scenarios

Data downtime and incident metrics

Mediumdata-quality-40
data-downtimereliability-engineeringslasmetrics

Question

How do you measure the reliability of your data, like uptime for services?

Solution

Data downtime measures the total duration that analytical data is missing, inaccurate, or unavailable for business operations, providing an objective reliability metric analogous to service outage downtime in web systems. It is calculated by multiplying the total number of data incidents by the sum of mean time to detect and mean time to resolve.

Quantifying pipeline downtime and mean time to recovery

Teams quantify platform health using a standardized downtime calculation:

  • The data downtime formula:
Data Downtime = Incident Count x (Mean Time to Detect + Mean Time to Resolve)

Reducing data downtime requires two complementary efforts: improving automated observability to detect issues immediately, and refining runbooks to accelerate recovery.

  • Mean Time to Detect (MTTD): Measures how long corrupted data or late batches sit in production before an engineer or monitor notices. Catching an anomaly via automated alerts in five minutes prevents days of silent metric corruption.
  • Mean Time to Resolve (MTTR): Measures how long it takes to diagnose the root cause, patch pipeline code, reprocess historical partitions, and restore production tables.

Key operational reliability metrics

Tracking comprehensive reliability metrics gives engineering leadership clear visibility into platform stability:

  • Freshness SLA misses: Track how frequently scheduled tables miss their delivery deadlines or exceed acceptable latency thresholds.
  • Automated test failure frequency: Measure the volume of schema, uniqueness, and referential integrity test failures across daily production runs.
  • Incident frequency and severity: Track monthly incident trends, categorizing events into high-severity business blockers and low-severity warnings.
  • Dataset-specific SLOs: Establish Service Level Objectives for tier-one tables, such as requiring daily financial transaction marts to be 99.9 percent complete and available by 7 AM UTC.
  • Blameless post-incident reviews: Conduct postmortems after every major outage to analyze root causes, update data contracts, and introduce automated regression assertions so similar failures cannot repeat.
PreviousNext