Databricks Asset Bundles let you describe jobs, pipelines, clusters and other resources in YAML files kept in Git, and deploy them to different environments from the command line. It is infrastructure as code for Databricks projects.
What a bundle is
A bundle is a project folder with a databricks.yml file plus source code and tests. The YAML describes resources and the targets they deploy to:
bundle:
name: orders_pipeline
resources:
jobs:
daily_orders:
name: daily_orders
tasks:
- task_key: ingest
notebook_task:
notebook_path: ./src/ingest.py
targets:
dev:
mode: development
workspace:
host: https://dev.cloud.databricks.com
prod:
mode: production
workspace:
host: https://prod.cloud.databricks.com
run_as:
service_principal_name: orders-prod-spCommands
databricks bundle validate databricks bundle deploy -t dev databricks bundle run daily_orders -t dev
A typical CI/CD flow
- Developers work in Git folders in the workspace, or locally in an IDE, on a feature branch.
- A pull request triggers CI: lint, unit tests (outside notebooks, against local Spark or Databricks Connect), and
bundle validate. - Merging to main deploys to a staging target, runs integration tests on sample data, and after approval deploys to production.
- Deployments run as a service principal, not a person, so production jobs do not depend on someone's account.
Practices that help
- Keep logic in Python modules or SQL files with tests, and keep notebooks thin. Business logic in notebooks cannot be tested well.
- Use parameters and targets for catalogs and paths, so the same code runs in dev and prod pointing at different data.
- Do not edit production jobs by hand in the UI. If you do, the next deploy overwrites it, or drift appears.
Say it in one line: bundles make Databricks jobs reviewable, repeatable and promotable across environments, just as any other code.