dbt (data build tool) is a transformation framework. It sits on top of your warehouse or lakehouse and turns raw tables into clean, tested, documented analytics models using SQL (and optionally Python).
The problem first
After ELT loads raw data into Snowflake, BigQuery, Redshift, Databricks, or similar, someone still has to:
- Rename and cast columns
- Join dimensions
- Build facts and marts
- Keep logic versioned, reviewed, and testable
Without dbt, that work often ends up as ad-hoc SQL scripts, notebooks, or BI calculated fields that nobody can CI-test or peer-review cleanly.
Mental model
Extract / Load (Airflow, Fivetran, custom jobs)
|
v
raw / landing tables in the warehouse
|
v
dbt models (SQL transforms + tests + docs)
|
v
marts / metrics tables for BI and appsWhat dbt does
1. You write SELECT models as .sql files (with Jinja). 2. dbt compiles them (resolving ref(), source(), macros). 3. dbt runs the compiled SQL in the warehouse as views, tables, merges, etc. 4. You attach schema tests, docs, and a dependency DAG.
What dbt is not
dbt does not extract from APIs or load CSV files into the warehouse by itself. Orchestrators and EL tools handle extract/load. dbt owns the T in ELT: transform data that is already in the warehouse.
Interview tip: Say "dbt brings software engineering practices to warehouse SQL: version control, testing, documentation, environments, and a dependency graph."