Learn · quality
Catching bad data before it reaches a dashboard someone trusts.
Why data quality matters, how to catch bad data before it reaches dashboards, quality gates, lineage, SLAs, and a publish-suite capstone.
Core foundational modules are free. Advanced production modules need Pro.
Playable walkthroughs: watch the system move, predict the next step, stamp a memory seal, then practice. Completing a Trace counts toward readiness.
Your pipeline succeeded. Every task is green. The dashboard shows revenue down 40 percent, and three hours of a leadership meeting go into working out whether the business is failing. It is not: an upstream team changed a currency field from dollars to cents. Nothing errored, because nothing was checking. A green pipeline is not the same as correct data, and the gap between those two is where trust is lost.
This is the difference between a pipeline that exists and a pipeline the business relies on. Being the person who catches a bad load before the CFO sees it is how data engineers build reputation, and quality work is increasingly its own job title.
A pipeline you have built
Quality checks need something to check. Any of the earlier tracks' capstones is enough.
SQL or pandas
Checks here run as real queries and frames, so basic fluency in one of them helps.
6 stages, in the order they build on each other. Each stage lists the modules and lessons it covers, and what you should be able to do by the end of it.
Quality is not a vibe. It breaks down into properties you can check: freshness, completeness, uniqueness, validity, and consistency. Naming them is what makes them testable.
What is Data QualityFree
0/2
Why bad data is expensive, what 'good data' means, and how testing strategies catch problems before dashboards do.
By the end of this stage
You can turn a vague complaint like 'the numbers look wrong' into a specific check.
The important design decision is where a check sits. A check that runs after publishing tells you that you already shipped the problem.
Quality GatesFree
0/3
Assert before you publish. Great-Expectations-shaped dicts, run on the real warehouse frames.
By the end of this stage
You can put a gate between transform and publish, and decide whether a failure should block the load or just warn.
Where people get stuck
Blocking on everything means someone disables the gates within a month. Decide which failures are worth stopping a pipeline for.
Row counts, null rates, ranges, referential integrity, and distribution drift each catch a different class of failure. Knowing which catches which is the skill.
More checks
0/3
Row counts, accepted statuses, and numeric bounds. Still pandas on df_orders.
By the end of this stage
You can pick the smallest set of checks that would have caught your last three incidents.
One check is a test. A suite plus lineage is an answer to the real question during an incident: what is affected downstream, and who do I need to tell?
Suites and lineage
0/2
Bundle checks into a job-level pass/fail, then name the columns that feed GMV.
By the end of this stage
You can trace an upstream break to the dashboards it affects.
Turn 'the data should be fresh' into a number with a consequence, because a promise without a number cannot be measured or defended.
SLA and SLI
0/1
The promise versus the measurement. The catalog would store both. You compute a ratio.
By the end of this stage
You can define a freshness SLA and the indicator that tells you when it is breached.
A publish suite that gates a real load, with alerts that a human would actually act on.
Quality capstone
0/1
One publish suite on df_orders: not-null, unique, row count. Print a summary whose flags are all True.
By the end of this stage
You can describe a quality strategy for a pipeline rather than a list of tests.
Finishing the lessons is not the goal. These are the things you should be able to do afterwards, and each one is worth checking honestly.
30 minutes a day
about 5 sessions
1 hour a day
about 3 sessions
4 hours a weekend day
about 1 session
Work through this track against a pipeline you already built in another track. Abstract quality lessons slide off. Adding a gate to your own capstone and watching it catch a deliberately broken row is what makes it stick.
These counts cover reading and the built-in exercises only. Real practice on the drills and a capstone will add to it, and that time is where most of the learning happens.
| Common mistake | What to do instead |
|---|---|
| Alerting on everything. | Alert fatigue means the real alert gets ignored. Page on things that need a human tonight, log the rest. |
| Checking only after the data is published. | That is detection, not prevention. Gate before the consumer sees it where you can. |
| Testing row counts and calling it quality. | The right number of wrong rows still passes. Check values, ranges, and relationships too. |