The dbt DAG (Directed Acyclic Graph) is the graph of models, sources, seeds, snapshots, and tests connected by dependencies. Nodes are resources; edges come mainly from ref(), source(), and test relationships.
Mental model
source(raw.orders) source(raw.customers)
| |
v v
stg_orders stg_customers
\ /
\ /
v v
fct_orders
|
v
dim_customer_orders (example mart)"Directed" means dependencies have a direction (upstream → downstream). "Acyclic" means no cycles: A cannot ref B if B already depends on A.
How dbt builds it
1. Parse all models and YAML. 2. Resolve every ref('x') / source('a','b'). 3. Order execution so upstream models run before downstream ones. 4. Power selectors like dbt run --select +fct_orders+ (parents, itself, children).
Why it matters
- Correct run order without manual scripts
- Impact analysis: "if I change
stg_orders, what breaks?" - Selective runs and CI on changed subgraphs
- Lineage in
dbt docs
Interview tip: Contrast with Airflow DAGs: Airflow orchestrates tasks/jobs; the dbt DAG describes data dependencies among SQL models inside the warehouse transform layer.