Databricks has a built-in orchestrator, Workflows, now called Lakeflow Jobs. A job is a set of tasks with dependencies between them, run on a schedule or by a trigger, on compute you choose for each task.
What a job can contain
Tasks can be notebooks, Python scripts or wheels, SQL queries, dbt projects, JARs, and declarative pipelines. You connect them in a graph, where a task starts after the tasks it depends on succeed. There are also conditional tasks (if/else), loops over a list of inputs (for-each), and the ability to run another job.
Triggers
- Schedule: a cron expression with a timezone.
- File arrival: start the job when new files land in a Unity Catalog location.
- Table update: start when an upstream table changes.
- Continuous: restart immediately after each run, for near-streaming workloads.
- API or manual runs.
Operating jobs
- Retries per task, timeouts, and notifications to email, Slack or webhooks.
- Repair runs: when a task fails in a run of ten, you fix the cause and repair, which reruns only the failed task and the tasks downstream of it, and does not redo the successful ones. That saves time and money.
- Each task can use its own job cluster or serverless compute, sized for that step.
- Run history, duration trends and parameters are available in the UI and through system tables.
Workflows or Airflow
Use Workflows when your pipeline lives mostly inside Databricks. It is integrated, has no separate system to run, and is simple to configure. Keep Airflow (or another orchestrator) when you orchestrate across many systems, such as loading from SaaS APIs, running dbt on another warehouse, calling services and then Databricks, or when the team already has mature Airflow operations. A common setup is Airflow triggering a Databricks job, and Databricks handling the tasks inside it.
Definitions as code
Define jobs in Databricks Asset Bundles (see the CI/CD question), so they are versioned and deployed the same way in each environment.