Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Where to put data quality checks

Data quality · Building Quality In

Where to put data quality checks

Easydata-quality-21
pipeline-stagesingestion-checksdata-qualitywrite-audit-publish

Question

At which points in a pipeline should data quality checks run?

Solution

Data quality checks belong at three distinct pipeline stages: during raw ingestion, immediately after business transformations, and right before publishing datasets to end consumers. Placing validation gates across these three boundaries catches bad records close to their origin and stops corrupted state from contaminating production reporting.

Where tests belong in the pipeline

Every layer of your lakehouse architecture requires targeted validations matching its responsibility:

  • Ingestion stage: Run cheap structural assertions as data lands in the raw zone. Verify schema types, reject unexpected column modifications, check for nulls in primary or partition keys, and flag batch volume anomalies such as empty files or unexpected record spikes.
  • Transformation stage: Test business logic on cleansed silver and gold tables. Check primary key uniqueness, foreign key referential integrity against dimension models, and domain constraints like positive order totals or valid transaction status codes.
  • Pre-publish stage: Verify operational health before updating consumer tables. Validate data freshness against arrival SLAs and reconcile aggregated metrics, such as comparing daily order counts and financial sums between the warehouse and upstream transaction logs.
Ingestion (Bronze)    -> Schema conformance, volume anomalies, null keys
Transformation (Silver)-> Uniqueness, referential integrity, business logic
Pre-publish (Gold)     -> Freshness SLAs, source-to-target reconciliation

The write-audit-publish workflow implements this progression cleanly:

  • Write staged output into an isolated directory or temporary table branch.
  • Audit the staged dataset using automated assertion suites.
  • Publish valid records to production only when every required test passes.

Trade-offs between run cost and coverage

Running every possible test on every micro-batch is too slow and expensive. Full-table uniqueness scans and cross-table referential joins on multi-billion row tables consume massive compute credits and delay batch completion.

Teams balance run costs by tiering assertions:

  • Run fast metadata and column checks on every pipeline execution.
  • Reserve heavy distribution drift analyses and full-table anti-joins for hourly or nightly reconciliations.
PreviousNext