Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. CI/CD with Databricks Asset Bundles

Snowflake, BigQuery & Databricks · Databricks

CI/CD with Databricks Asset Bundles

Mediumwarehouses-51
databricksasset-bundlesci-cddeploymentgit

Question

How do you do CI/CD for Databricks?

Solution

Databricks Asset Bundles let you describe jobs, pipelines, clusters and other resources in YAML files kept in Git, and deploy them to different environments from the command line. It is infrastructure as code for Databricks projects.

What a bundle is

A bundle is a project folder with a databricks.yml file plus source code and tests. The YAML describes resources and the targets they deploy to:

bundle:
  name: orders_pipeline

resources:
  jobs:
    daily_orders:
      name: daily_orders
      tasks:
        - task_key: ingest
          notebook_task:
            notebook_path: ./src/ingest.py

targets:
  dev:
    mode: development
    workspace:
      host: https://dev.cloud.databricks.com
  prod:
    mode: production
    workspace:
      host: https://prod.cloud.databricks.com
    run_as:
      service_principal_name: orders-prod-sp

Commands

databricks bundle validate
databricks bundle deploy -t dev
databricks bundle run daily_orders -t dev

A typical CI/CD flow

  • Developers work in Git folders in the workspace, or locally in an IDE, on a feature branch.
  • A pull request triggers CI: lint, unit tests (outside notebooks, against local Spark or Databricks Connect), and bundle validate.
  • Merging to main deploys to a staging target, runs integration tests on sample data, and after approval deploys to production.
  • Deployments run as a service principal, not a person, so production jobs do not depend on someone's account.

Practices that help

  • Keep logic in Python modules or SQL files with tests, and keep notebooks thin. Business logic in notebooks cannot be tested well.
  • Use parameters and targets for catalogs and paths, so the same code runs in dev and prod pointing at different data.
  • Do not edit production jobs by hand in the UI. If you do, the next deploy overwrites it, or drift appears.

Say it in one line: bundles make Databricks jobs reviewable, repeatable and promotable across environments, just as any other code.

PreviousNext