An expectation is a data quality rule attached to a dataset in a declarative pipeline. Each record is checked against a SQL condition, and you choose what happens when it fails: record the failure and keep the row, drop the row, or fail the update.
Three actions
CREATE OR REFRESH STREAMING TABLE silver_orders ( CONSTRAINT valid_id EXPECT (order_id IS NOT NULL), CONSTRAINT valid_amount EXPECT (amount >= 0) ON VIOLATION DROP ROW, CONSTRAINT valid_date EXPECT (order_date >= '2020-01-01') ON VIOLATION FAIL UPDATE ) AS SELECT * FROM STREAM bronze_orders;
What each action does:
EXPECTalone: the invalid rows are kept in the table, and the failures are counted in the metrics. Use it for rules you want to monitor without blocking.ON VIOLATION DROP ROW: invalid rows are removed from the target, and counted.ON VIOLATION FAIL UPDATE: the whole pipeline update stops at the first failing record. Use this for rules where bad data must never get through.
In Python, the same options are decorators, such as expect, expect_or_drop and expect_or_fail.
Metrics
For each expectation, the pipeline's event log records how many records passed and failed in each update. You can query the event log, chart the failure rate and alert on it, so quality becomes something you watch over time.
Quarantine pattern
Dropping bad rows silently is dangerous, because you lose the evidence. A common pattern is to send them somewhere else: write a second table with the inverse condition.
CREATE OR REFRESH STREAMING TABLE quarantine_orders AS SELECT * FROM STREAM bronze_orders WHERE NOT (order_id IS NOT NULL AND amount >= 0);
People can then inspect them, fix upstream causes, and reprocess.
Choosing the strictness
Apply light checks on bronze (is it parseable), stricter checks on silver, and fail only for rules that protect downstream consumers, such as primary keys on financial data. Too many fail-update rules make pipelines fragile. Too few let junk through. Tune them based on the failure metrics.