Data quality checks should block a pipeline when bad data causes irreversible downstream harm, such as corrupted financial numbers or broken joins. Checks should only issue warnings on soft signals like minor volume fluctuations or subtle statistical shifts that warrant investigation without halting operational pipelines.
Deciding between hard fails and soft alerts
Deciding whether a test failure halts a pipeline depends on the business impact of letting bad data reach production:
- Block execution: Halt the orchestrator when bad records corrupt downstream tables. Broken schemas, duplicate primary keys in financial ledgers, and null foreign keys in join paths must block publishing immediately. Allowing duplicated transaction IDs into a revenue mart produces inflated reports that misinform executive leadership.
- Warn and proceed: Emit an alert but allow the pipeline to complete when signals reflect potential drift rather than hard invalidity. A daily batch that is ten percent smaller than normal or a slight shift in customer category distributions requires investigation, but stopping the entire company dashboard over minor variance damages user trust.
Hard Fail (Block) -> Duplicate primary keys, null join keys, schema mismatch Soft Alert (Warn) -> 10% volume drop, unexpected category drift, latency bump
Managing these checks effectively requires establishing clear operational practices:
- Assign explicit severity levels in your orchestration DAGs, such as critical errors that fail task execution and informational warnings that send asynchronous Slack notifications.
- Prevent alert fatigue by silencing redundant warnings and only routing actionable emergencies to on-call engineers.
- Periodically review check thresholds as business volume grows, adjusting static limits to prevent stale assumptions from triggering false alarms.
Operational stability and team focus
When pipelines fail on every minor warning, teams become desensitized to alerts and start ignoring real emergencies. Reserving hard blocks for true data corruption guarantees that on-call engineers treat every pipeline halt as an urgent, genuine incident.