Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Data Quality & Observability

Progress0/12
x

What is Data Quality

  • What is data quality?12m
  • Testing strategies for data12m

Quality Gates

  • Why quality gates12m
  • Not-null on keys and measures12m
  • Uniqueness on the grain12m

More checks

  • Row count between bounds10m
  • Accepted values12m
  • Freshness and numeric bounds12m

Suites and lineage

  • Run a suite14m
  • Data lineage12m

SLA and SLI

  • SLA versus SLI12m

Quality capstone

  • Capstone: a publish suite16m
Back to track
  1. Learn
  2. Data Quality & Observability
  3. What is Data Quality
  4. Testing strategies for data

Lesson 2 of 12 · Theory first, then run it

Testing strategies for data

qualitypandasbeginner12 min

Overview

Unit, integration, and contract tests catch different failures. A tiny suite of two checks prints whether both not-null and unique passed.

On this page7 sections›
  1. 1The idea
  2. 2Why this exists
  3. 3Picture this
  4. 4A small example
  5. 5Common beginner questions
  6. 6What comes next
  7. 7Practice

The idea

A unit check tests one column or one helper in isolation: is order_id non-null, is order_id unique. An integration check tests a transform that combines frames: does every payment join to an order. A contract check freezes the shape a consumer depends on: columns present, types stable, statuses in a named set.

Quality gates, dbt tests, and assertions are three envelopes for those checks. A gate returns a dict you can log, screenshot, and fan into a suite. A dbt test is a SQL query that fails when it returns rows. An assertion throws and stops the process.

This lesson stays on one frame so the suite is tiny: two unit checks on df_orders.order_id. Later lessons add row-count, accepted values, and a job-level all_passed bit. The instinct is the same at every scale: name the checks, run them, print whether they all passed.

Why this exists

An e-commerce checkout path can pass a unit test on order_id while an inner join to a broken payments extract drops every row. Gold is empty. The unit test is still green. Finance opens a revenue dashboard of zeros and treats it as a quiet day. Integration checks (row count after the join, orphan payments) catch that class of failure.

A bank's fraud dashboard depends on a status column staying inside a closed set. A vendor ships a new token, PENDING_REVIEW. The unit not-null check still passes. The chart grows a bar the mapping never handled. A contract test on accepted values would have failed the publish.

Testing strategy is choosing which envelope and which layer. You do not pick one forever. You run unit checks on keys, integration checks on joins, and contract checks on the columns a dashboard or a downstream job already coded against.

Picture this

Three layers of data tests
UnitIntegration and contractlayerUnit: not-null order_idUnit: unique order_idFast, cheap, localIntegration: payments join ordersContract: status in allowed setCatches empty gold and new tokens

Unit is one column. Integration is the join. Contract is the consumer's shape.

A quality gate is a function that returns {"success": bool, "failed_rows": int} (or a close cousin such as duplicate_count). You print the dict. You still assign a DataFrame to result. The next module teaches that helper shape. Tonight the suite is a dict of booleans so you can see all() reduce them.

Same instinct. Different place in the stack. This tab uses pandas gates on warehouse frames.

EnvelopeWhat a failure looks likeWhen to use it
Assertion (assert)Exception, process diesLocal notebooks, schema-contracts
Quality gate (dict)Printed {success, failed_rows}Jobs you want to log and fan into a suite
dbt test (SQL)Query returns unexpected rowsWarehouse models after dbt build
Contract (shape + set)Missing column or new tokenAnything a dashboard already mapped

Build the suite as a dict so a reviewer can read the keys. all(suite.values()) is True only when every check passed. Print that boolean. Later, qa-suite will collect expectation dicts and reduce on item["success"] instead of bare booleans.

  1. Pick the grain column (order_id on df_orders).
  2. Unit: count nulls. Unit: count extras from duplicated().
  3. Store each as a boolean in a suite dict.
  4. Print all(suite.values()). Project the grain column into result.

df_orders.order_id passes both checks in this warehouse. df_events.event_id does not pass unique: late ingest replays a slice of events. A unique unit check on the wrong grain is a false alarm, not a broken pipeline.

  • Unit checks are cheap. Run them on every key and measure you will invoice from.
  • Integration checks need two frames. Row count after a join is the simplest one.
  • Contract checks name the set and the columns. Do not build the set from unique() of today's data.

Print whether all passed. Do not assign the boolean to result. The grid needs a DataFrame.

PythonTwo unit checks, one printed bit
suite = {
    "not_null": int(df_orders["order_id"].isna().sum()) == 0,
    "unique": int(df_orders["order_id"].duplicated().sum()) == 0,
}
print(all(suite.values()))
result = df_orders.loc[:, ["order_id"]]

The same two helpers on events show a split: not-null may pass while unique fails. That is why a suite prints one bit and still keeps the per-check keys for the incident channel.

PythonPassing grain versus a grain that is allowed to repeat
orders_suite = {
    "not_null": int(df_orders["order_id"].isna().sum()) == 0,
    "unique": int(df_orders["order_id"].duplicated().sum()) == 0,
}
events_suite = {
    "not_null": int(df_events["event_id"].isna().sum()) == 0,
    "unique": int(df_events["event_id"].duplicated().sum()) == 0,
}
print("orders", all(orders_suite.values()), orders_suite)
print("events", all(events_suite.values()), events_suite)
result = df_orders.loc[:, ["order_id"]]

Do not unique() your way to a contract

allowed = set(df_orders["order_status"].unique()) always passes. That tests the function, not the consumer. Name the set the dashboard already coded against.

Orchestration green is not a test suite

A scheduler reports that the Python process exited 0. That is not completeness, uniqueness, or a contract. The Orchestration track taught DAGs and retries. This track teaches the checks you run before the publish task is allowed to succeed.

A small example

Build a tiny suite on df_orders.order_id: one not-null check and one unique check. Print whether all passed. The per-check keys stay in the dict for the log.

Input: two unit checks on the order grain.

checkwhat it testsexpected here
not_nullisna().sum() == 0pass
uniqueduplicated().sum() == 0pass
PythonTwo checks, one printed bit
suite = {
    "not_null": int(df_orders["order_id"].isna().sum()) == 0,
    "unique": int(df_orders["order_id"].duplicated().sum()) == 0,
}
print(all(suite.values()))
print(suite)
result = df_orders.loc[:, ["order_id"]]

stdout shows True because both checks pass on df_orders. df_events.event_id would print False on unique because late ingest replays a slice.

Output: the suite bit is True only when every check passes.

frameall passed?why
df_orders.order_idTrueclean key at the claimed grain
df_events.event_idFalselate replay allowed on events

Copy-paste without reading the output

Run Sample first. If the numbers or row count look wrong, stop and re-read the previous section before changing code.

Common beginner questions

Is a quality gate the same as a dbt test?

Same instinct, different runtime. dbt tests are SQL in a warehouse project. Quality gates here are pandas functions that return a dict. Great Expectations calls them expectations. You would swap these helpers for that library on a laptop.

Should every check throw?

Assertions are right when a local notebook must not continue. Gates are right when a job should log every failure, then decide. Suites need the dict shape. schema-contracts remains the assert cousin.

Where do integration tests live if this lesson only uses df_orders?

On the join. A later row-count gate is the cheap version: after the transform, the frame must not be empty. Full join tests belong next to the transform that produced the join.

What comes next

The next lesson turns a completeness count into a quality gate: a function that returns success and failed_rows, printed as a dict. That envelope is what you will reuse for the rest of the track.

Practice

Run Sample to see both checks pass on df_orders. Then complete Exercise: build a suite dict with not-null and unique on df_orders["order_id"], print whether all passed, and assign result = df_orders.loc[:, ["order_id"]].

stdout should include True. Use isna for the null check and duplicated for the unique check.

Practicals · load into the editor

After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.

Practice this

Same ideas as interview drills. These challenges open in the studio with a dataset and tests already set up.

  • Quality gate: telemetry rows missing a readingProduction ticket: expect_column_not_null on telemetry.reading - return the actual violating rows, not just a pass/fail.Studiobeginnerpandas12 min
Rate:
Was this useful?
What is data quality?Why quality gates