Overview
A green DAG that wrote null order_ids is not a successful night. The gate is a function that returns success and a count, then you print it.
On this page7 sections
What happened
A quality gate is a check that runs after the transform and before you publish gold. It answers a yes-or-no question about the frame: are there null keys, duplicate grains, an empty table. The scheduler's green light only means the process exited 0.
You will write gates as small pandas functions that return {"success": bool, "failed_rows": int}. Print the dict so logs and graders can see it. Assign a column projection to result so the grid has a DataFrame. Do not put the dict in result.
Great Expectations calls these checks expectations. dbt calls them tests. The pandas schema-contracts lesson used assert, which throws. This track reports, so a suite can keep going past the first failure and still print every count.
Why it broke (or worked)
An e-commerce gold table of GMV can load on time with null order_ids in a slice of rows. The DAG is green. Finance invoices from the table. Refunds cannot find those rows. A not-null gate on order_id would have stopped the publish and left yesterday's gold in place.
A bank's nightly transfer mart can finish while a join to accounts dropped every unmatched id. The job exited 0 because Python never raised. The mart is empty. A gate that counts failed rows (or a row-count bound, later in this track) is the difference between a quiet success and a page.
Gates exist because dashboards and schedulers do not inspect grain. They render and they exit. The check has to sit in the path between transform and publish, on the real warehouse frames, not on a mock.
The pattern to steal
The gate sits after the frame exists and before gold is trusted.
Count isna() on the column, store the count in failed_rows, and set success when the count is zero. int() the numpy count so the printed dict looks like a Python dict a log shipper will not mangle.
Same warehouse. Three envelopes for the same instinct.
| Style | What it does | Where you saw it |
|---|---|---|
| assert col.is_unique | Throws, job dies | Pandas schema-contracts |
| SELECT ... HAVING COUNT(*) > 1 | Return violators; empty is pass | SQL data-quality-constraints, dbt unique |
| {"success", "failed_rows"} | Printable, composable | This track |
The tables are real warehouse frames: df_orders, df_events, df_order_items, df_payments, df_customers, df_telemetry, df_customer_cdc. Sample also runs the helper on df_telemetry.reading so you can see a False when sensor rows arrive blank.
- Write expect_column_not_null(df, col).
- failed_rows = int(df[col].isna().sum()).
- Return {"success": failed_rows == 0, "failed_rows": failed_rows}.
- Print the dict. Project a column list into result.
Printing is part of the contract. A gate that returns a dict you never print is invisible in logs. The exercise looks at stdout for True because order_id has no nulls here.
- Return a dict. Print it. Put a DataFrame in result.
- schema-contracts remains the assert cousin. Do not rewrite that lesson's dropna prompt here.
- failed_rows is a count of violators, not a list of them. Suites reduce on success.
Call the helper on a clean column and on a dirty one so you trust both branches before the exercise grades the clean path.
def expect_column_not_null(df, col):
failed_rows = int(df[col].isna().sum())
return {"success": failed_rows == 0, "failed_rows": failed_rows}
print(expect_column_not_null(df_orders, "order_id"))
result = df_orders.loc[:, ["order_id"]]Telemetry is the failing lane. A False in the dict is not a broken helper. It is the helper doing its job on a column that is allowed to have gaps.
def expect_column_not_null(df, col):
failed_rows = int(df[col].isna().sum())
return {"success": failed_rows == 0, "failed_rows": failed_rows}
print("orders.order_id", expect_column_not_null(df_orders, "order_id"))
print("telemetry.reading", expect_column_not_null(df_telemetry, "reading"))
result = df_orders.loc[:, ["order_id"]]Do not assign the dict to result
The pandas runner errors if result is a dict. Print the expectation. Project columns into result. The grid displays frames, not gate payloads.
A green DAG is an exit code
The Orchestration track taught retries, idempotent loads, and DAGs. Exit 0 means the process finished. It does not mean order_id is non-null. Put the gate in the publish path of that DAG.
How you write it
Turn a completeness count into a gate dict you can print and fan into a suite. Start with order_id (clean) and telemetry.reading (gappy).
Input: two orders with filled keys.
| order_id | order_total |
|---|---|
| ORD-0001 | 84.20 |
| ORD-0002 | 31.00 |
def expect_column_not_null(df, col):
failed_rows = int(df[col].isna().sum())
return {"success": failed_rows == 0, "failed_rows": failed_rows}
print(expect_column_not_null(df_orders, "order_id"))
print(expect_column_not_null(df_telemetry, "reading"))
result = df_orders.loc[:, ["order_id"]]The orders call prints success True and failed_rows 0. The telemetry call prints success False with a positive failed_rows count.
Output: the dict shape you will reuse for every gate in this track.
| column | success | failed_rows |
|---|---|---|
| order_id | True | 0 |
| reading | False | > 0 |
Copy-paste without reading the output
Run Sample first. If the numbers or row count look wrong, stop and re-read the previous section before changing code.
Common beginner questions
Why not use assert like schema-contracts?
assert stops at the first failure and has nothing to log besides a traceback. A dict lets you run three checks, print three counts, and still fail the job on all_passed. You will build that suite later.
Why print if you already return?
Logs and graders read stdout. Return values vanish unless a caller stores them. Print the dict every time you run a gate in this tab.
Is Great Expectations required?
No. The library is the production shape of these helpers. This tab teaches the dict so you can defend the instinct. On a laptop you would swap the helpers for that library, not invent a new one.
What comes next
The next lesson runs the same not-null gate twice: once on the key (order_id) and once on the measure (order_total). One function, two stories.
Practice
Run Sample to see True on order_id and False on telemetry.reading. Then complete Exercise: implement expect_column_not_null(df, col) returning {"success": bool, "failed_rows": int}. Call it on df_orders, "order_id" and print the dict. Assign result = df_orders.loc[:, ["order_id"]].
stdout should include True. Define and call expect_column_not_null. failed_rows = int(df[col].isna().sum()).
Practicals · load into the editor
After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.