Data teams run dbt from orchestrators like Apache Airflow using three common patterns: executing CLI commands with bash or container operators, rendering models as individual tasks using Astronomer Cosmos, or triggering remote jobs through dbt Cloud API operators. Each approach balances setup simplicity against granular task monitoring and retry capabilities.
Coarse execution with bash operators
The simplest method runs dbt as a single monolithic task using BashOperator or DockerOperator:
run_dbt = BashOperator(
task_id="dbt_build_daily",
bash_command="cd /opt/dbt && dbt build --target prod",
)This pattern is easy to implement and isolates dbt dependencies cleanly inside containers. However, it treats the entire dbt project as a black box:
- Airflow cannot see individual model status in its user interface.
- If model number ninety-eight fails out of one hundred, retrying the Airflow task re-runs the entire pipeline from the beginning unless you manually write custom selection flags.
This approach works best for small projects with short total runtimes.
Astronomer Cosmos and fine-grained tasks
Astronomer Cosmos parses the dbt project manifest.json and automatically converts each dbt model, snapshot, and test into a native Airflow task:
- The Airflow UI displays the full dbt lineage graph visually inside the DAG view.
- If a single mart model fails, an engineer can retry just that failed model and its downstream dependents directly from the Airflow UI without restarting upstream tasks.
- Task execution can use Airflow pool limits and worker distribution across Celery or Kubernetes workers.
The tradeoff is scheduler overhead: rendering hundreds of dbt models as individual Airflow tasks increases metadata database queries and DAG parsing times.
Cloud API triggers and execution metadata
When using dbt Cloud or dbt platform, the recommended integration uses the DbtCloudRunJobOperator:
- Airflow sends an API request to trigger a pre-configured job in dbt Cloud, then polls for completion.
- Transformation compute, logging, and artifact generation run entirely within dbt Cloud, offloading resource pressure from the Airflow cluster.
- You can pass Airflow logical execution dates into dbt models as variables, such as passing execution_date with the Airflow ds macro.
- After completion, the pipeline downloads run_results.json and pushes it to object storage for audit logging and observability tools.
This pattern combines centralized orchestration with dedicated transformation management.