Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Data Quality & Observability

Progress0/12
x

What is Data Quality

  • What is data quality?12m
  • Testing strategies for data12m

Quality Gates

  • Why quality gates12m
  • Not-null on keys and measures12m
  • Uniqueness on the grain12m

More checks

  • Row count between bounds10m
  • Accepted values12m
  • Freshness and numeric bounds12m

Suites and lineage

  • Run a suite14m
  • Data lineage12m

SLA and SLI

  • SLA versus SLI12m

Quality capstone

  • Capstone: a publish suite16m
Back to track
  1. Learn
  2. Data Quality & Observability
  3. Quality Gates
  4. Why quality gates

Lesson 3 of 12 · Theory first, then run it

Why quality gates

qualitypandasbeginner12 min

Overview

A green DAG that wrote null order_ids is not a successful night. The gate is a function that returns success and a count, then you print it.

On this page7 sections›
  1. 1What happened
  2. 2Why it broke (or worked)
  3. 3The pattern to steal
  4. 4How you write it
  5. 5Common beginner questions
  6. 6What comes next
  7. 7Practice

What happened

A quality gate is a check that runs after the transform and before you publish gold. It answers a yes-or-no question about the frame: are there null keys, duplicate grains, an empty table. The scheduler's green light only means the process exited 0.

You will write gates as small pandas functions that return {"success": bool, "failed_rows": int}. Print the dict so logs and graders can see it. Assign a column projection to result so the grid has a DataFrame. Do not put the dict in result.

Great Expectations calls these checks expectations. dbt calls them tests. The pandas schema-contracts lesson used assert, which throws. This track reports, so a suite can keep going past the first failure and still print every count.

Why it broke (or worked)

An e-commerce gold table of GMV can load on time with null order_ids in a slice of rows. The DAG is green. Finance invoices from the table. Refunds cannot find those rows. A not-null gate on order_id would have stopped the publish and left yesterday's gold in place.

A bank's nightly transfer mart can finish while a join to accounts dropped every unmatched id. The job exited 0 because Python never raised. The mart is empty. A gate that counts failed rows (or a row-count bound, later in this track) is the difference between a quiet success and a page.

Gates exist because dashboards and schedulers do not inspect grain. They render and they exit. The check has to sit in the path between transform and publish, on the real warehouse frames, not on a mock.

The pattern to steal

Transform, gate, publish
df_orders in memoryexpectation dictprint and logpublish projection

The gate sits after the frame exists and before gold is trusted.

Count isna() on the column, store the count in failed_rows, and set success when the count is zero. int() the numpy count so the printed dict looks like a Python dict a log shipper will not mangle.

Same warehouse. Three envelopes for the same instinct.

StyleWhat it doesWhere you saw it
assert col.is_uniqueThrows, job diesPandas schema-contracts
SELECT ... HAVING COUNT(*) > 1Return violators; empty is passSQL data-quality-constraints, dbt unique
{"success", "failed_rows"}Printable, composableThis track

The tables are real warehouse frames: df_orders, df_events, df_order_items, df_payments, df_customers, df_telemetry, df_customer_cdc. Sample also runs the helper on df_telemetry.reading so you can see a False when sensor rows arrive blank.

  1. Write expect_column_not_null(df, col).
  2. failed_rows = int(df[col].isna().sum()).
  3. Return {"success": failed_rows == 0, "failed_rows": failed_rows}.
  4. Print the dict. Project a column list into result.

Printing is part of the contract. A gate that returns a dict you never print is invisible in logs. The exercise looks at stdout for True because order_id has no nulls here.

  • Return a dict. Print it. Put a DataFrame in result.
  • schema-contracts remains the assert cousin. Do not rewrite that lesson's dropna prompt here.
  • failed_rows is a count of violators, not a list of them. Suites reduce on success.

Call the helper on a clean column and on a dirty one so you trust both branches before the exercise grades the clean path.

PythonPrint the gate, project the frame
def expect_column_not_null(df, col):
    failed_rows = int(df[col].isna().sum())
    return {"success": failed_rows == 0, "failed_rows": failed_rows}

print(expect_column_not_null(df_orders, "order_id"))
result = df_orders.loc[:, ["order_id"]]

Telemetry is the failing lane. A False in the dict is not a broken helper. It is the helper doing its job on a column that is allowed to have gaps.

PythonPass on the key, fail on the sensor
def expect_column_not_null(df, col):
    failed_rows = int(df[col].isna().sum())
    return {"success": failed_rows == 0, "failed_rows": failed_rows}

print("orders.order_id", expect_column_not_null(df_orders, "order_id"))
print("telemetry.reading", expect_column_not_null(df_telemetry, "reading"))
result = df_orders.loc[:, ["order_id"]]

Do not assign the dict to result

The pandas runner errors if result is a dict. Print the expectation. Project columns into result. The grid displays frames, not gate payloads.

A green DAG is an exit code

The Orchestration track taught retries, idempotent loads, and DAGs. Exit 0 means the process finished. It does not mean order_id is non-null. Put the gate in the publish path of that DAG.

How you write it

Turn a completeness count into a gate dict you can print and fan into a suite. Start with order_id (clean) and telemetry.reading (gappy).

Input: two orders with filled keys.

order_idorder_total
ORD-000184.20
ORD-000231.00
PythonGate on a clean key and a dirty sensor
def expect_column_not_null(df, col):
    failed_rows = int(df[col].isna().sum())
    return {"success": failed_rows == 0, "failed_rows": failed_rows}

print(expect_column_not_null(df_orders, "order_id"))
print(expect_column_not_null(df_telemetry, "reading"))
result = df_orders.loc[:, ["order_id"]]

The orders call prints success True and failed_rows 0. The telemetry call prints success False with a positive failed_rows count.

Output: the dict shape you will reuse for every gate in this track.

columnsuccessfailed_rows
order_idTrue0
readingFalse> 0

Copy-paste without reading the output

Run Sample first. If the numbers or row count look wrong, stop and re-read the previous section before changing code.

Common beginner questions

Why not use assert like schema-contracts?

assert stops at the first failure and has nothing to log besides a traceback. A dict lets you run three checks, print three counts, and still fail the job on all_passed. You will build that suite later.

Why print if you already return?

Logs and graders read stdout. Return values vanish unless a caller stores them. Print the dict every time you run a gate in this tab.

Is Great Expectations required?

No. The library is the production shape of these helpers. This tab teaches the dict so you can defend the instinct. On a laptop you would swap the helpers for that library, not invent a new one.

What comes next

The next lesson runs the same not-null gate twice: once on the key (order_id) and once on the measure (order_total). One function, two stories.

Practice

Run Sample to see True on order_id and False on telemetry.reading. Then complete Exercise: implement expect_column_not_null(df, col) returning {"success": bool, "failed_rows": int}. Call it on df_orders, "order_id" and print the dict. Assign result = df_orders.loc[:, ["order_id"]].

stdout should include True. Define and call expect_column_not_null. failed_rows = int(df[col].isna().sum()).

Practicals · load into the editor

After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.

Rate:
Was this useful?
Testing strategies for dataNot-null on keys and measures