Learn · dbt
How teams keep hundreds of SQL models organised, tested, and documented.
What analytics engineering is, how dbt compiles and tests SQL models, staging layers, marts, and incrementals. SQL runs in the editor; no dbt CLI needed.
Core foundational modules are free. Advanced production modules need Pro.
Playable walkthroughs: watch the system move, predict the next step, stamp a memory seal, then practice. Completing a Trace counts toward readiness.
A team of five analysts each write their own revenue query. Four of them are subtly different, two are wrong, and nobody knows which one the CFO's dashboard uses. Someone renames a column upstream and six dashboards break silently on Monday. The queries are all fine individually. The problem is that there is no shared, tested, dependency-aware place for them to live. That is the problem dbt solves.
dbt is now the default in the analytics engineering half of the field. The valuable skill is not the CLI, it is knowing how to layer models, where to put business logic, and what to test so a broken upstream change fails loudly instead of quietly.
dbt is SQL plus structure. If joins, CTEs, and aggregation are not comfortable yet, do the SQL track first.
Knowing what a fact and a dimension are makes the marts module land properly.
6 stages, in the order they build on each other. Each stage lists the modules and lessons it covers, and what you should be able to do by the end of it.
Before the tool, the role. There is a gap between the raw tables a data engineer lands and the numbers a business trusts, and this stage is about who fills it and how.
Analytics EngineeringFree
0/3
What analytics engineering is, how a dbt project is laid out, and how Jinja templates become SQL.
By the end of this stage
You can explain what analytics engineering is and why the role appeared.
dbt looks magical until you see that it compiles SQL, works out the dependency order, and runs it. Then it looks obvious, which is the right feeling.
What dbt IsFree
0/2
The SQL really runs. The dbt CLI around it does not. source() and ref() become relation names.
By the end of this stage
You can explain what happens between writing a model file and seeing a table in the warehouse.
Raw source tables are other teams' shapes, with their names and their quirks. Cleaning them once in a staging layer stops that mess spreading into every downstream model.
Staging
0/2
Rename, cast, and light-filter raw sources. Tests start at not_null.
By the end of this stage
You can build a staging model that renames, casts, and standardises a raw source cleanly.
The point of tests here is not coverage, it is early warning. A unique test on a key catches the duplicate that would have doubled revenue on a dashboard.
unique and relationships
0/2
Generic tests are SQL that must return zero rows. unique and relationships are the other two you will write on every staging model.
By the end of this stage
You can pick the tests that would have caught your last data incident, and explain why each one matters.
Where people get stuck
Testing everything is as unhelpful as testing nothing. Test the assumptions that would cause visible damage if they broke.
Rebuilding every model from scratch every hour stops being viable as data grows. Incremental models only process what is new, which introduces the same duplicate and late-data problems you met in orchestration.
Marts and Incremental
0/3
Business-shaped SELECTs on top of staging, then the incremental filter you would wrap in a materialization.
By the end of this stage
You can build a mart on a stated grain and make it incremental without double counting.
A layered project from raw source to tested mart, which is exactly what a dbt interview asks you to describe.
By the end of this stage
You can walk through a source to staging to mart lineage and defend each layer's job.
Finishing the lessons is not the goal. These are the things you should be able to do afterwards, and each one is worth checking honestly.
30 minutes a day
about 6 sessions
1 hour a day
about 3 sessions
4 hours a weekend day
about 1 session
SQL runs in the editor here, so you do not need the dbt CLI or a warehouse account to learn the concepts. If you have a local Postgres, running the real dbt CLI afterwards is worth an afternoon, because seeing compiled SQL and a lineage graph makes the model click.
These counts cover reading and the built-in exercises only. Real practice on the drills and a capstone will add to it, and that time is where most of the learning happens.
| Common mistake | What to do instead |
|---|---|
| Putting business logic in staging models. | Staging cleans and renames. Logic belongs downstream, or every consumer inherits a decision they did not ask for. |
| Adding tests to every column to feel thorough. | Noisy tests get ignored, and ignored tests are worse than no tests. Test the keys, the relationships, and the assumptions with real consequences. |
| Making a model incremental before it needs to be. | Incremental adds a class of subtle bugs. Do it when full refresh actually hurts, not by default. |