Data quality checks belong at three distinct pipeline stages: during raw ingestion, immediately after business transformations, and right before publishing datasets to end consumers. Placing validation gates across these three boundaries catches bad records close to their origin and stops corrupted state from contaminating production reporting.
Where tests belong in the pipeline
Every layer of your lakehouse architecture requires targeted validations matching its responsibility:
- Ingestion stage: Run cheap structural assertions as data lands in the raw zone. Verify schema types, reject unexpected column modifications, check for nulls in primary or partition keys, and flag batch volume anomalies such as empty files or unexpected record spikes.
- Transformation stage: Test business logic on cleansed silver and gold tables. Check primary key uniqueness, foreign key referential integrity against dimension models, and domain constraints like positive order totals or valid transaction status codes.
- Pre-publish stage: Verify operational health before updating consumer tables. Validate data freshness against arrival SLAs and reconcile aggregated metrics, such as comparing daily order counts and financial sums between the warehouse and upstream transaction logs.
Ingestion (Bronze) -> Schema conformance, volume anomalies, null keys Transformation (Silver)-> Uniqueness, referential integrity, business logic Pre-publish (Gold) -> Freshness SLAs, source-to-target reconciliation
The write-audit-publish workflow implements this progression cleanly:
- Write staged output into an isolated directory or temporary table branch.
- Audit the staged dataset using automated assertion suites.
- Publish valid records to production only when every required test passes.
Trade-offs between run cost and coverage
Running every possible test on every micro-batch is too slow and expensive. Full-table uniqueness scans and cross-table referential joins on multi-billion row tables consume massive compute credits and delay batch completion.
Teams balance run costs by tiering assertions:
- Run fast metadata and column checks on every pipeline execution.
- Reserve heavy distribution drift analyses and full-table anti-joins for hourly or nightly reconciliations.