Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. All-purpose vs job clusters vs serverless

Snowflake, BigQuery & Databricks · Databricks

All-purpose vs job clusters vs serverless

Easywarehouses-38
databrickscomputejob-clustersserverlesssql-warehouse

Question

What is the difference between all-purpose clusters, job clusters and serverless compute in Databricks?

Solution

Databricks offers several kinds of compute, and the main difference is who uses them and how long they live. All-purpose clusters are for interactive work. Job compute is created for one job run and then deleted. Serverless compute is managed by Databricks and starts quickly.

All-purpose clusters

You create them to work interactively in notebooks. Several people can attach to the same one, and it stays running until you stop it or auto-termination kicks in after an idle period. They are convenient and also the most expensive way to run a scheduled job, because they bill at a higher rate and often sit idle.

Job clusters

When a scheduled job runs, Databricks creates a fresh cluster for that run and terminates it as soon as the job ends. It has a lower rate than an all-purpose cluster, and nothing sits idle between runs. The cost is the startup time of a few minutes for each run, and no sharing between runs. For production pipelines, this is the usual default on classic compute.

Serverless compute

You do not choose machine types or cluster sizes. Databricks provisions the compute and starts it in seconds. It exists for notebooks, jobs and pipelines. You trade fine control (such as custom libraries at the OS level) for simplicity and fast startup. It suits spiky workloads, where waiting minutes for a cluster is annoying.

SQL warehouses

A separate type of compute optimised for SQL and BI queries, in classic, pro and serverless variants. Analysts and BI tools such as Power BI or Tableau connect to them. They come with the Photon engine and handle many concurrent queries, and they can scale out with more clusters.

Auto-termination

Set an idle timeout on every interactive cluster, such as 30 to 60 minutes, so a forgotten cluster does not run all weekend.

A good rule

Develop on an all-purpose cluster or serverless notebook, and run production on job compute or serverless jobs. Use SQL warehouses for dashboards. Mention that you would use cluster policies to enforce sensible sizes and auto-termination for everyone.

PreviousNext