Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Databricks architecture: control plane and compute plane

Snowflake, BigQuery & Databricks · Databricks

Databricks architecture: control plane and compute plane

Easywarehouses-37
databricksarchitecturecontrol-planecompute-plane

Question

Explain the Databricks architecture.

Solution

Databricks splits its platform into a control plane, which Databricks runs, and a compute plane, where your data is processed. Your data stays in your own cloud storage the whole time.

Control plane (Databricks account)         Compute plane
  web UI, notebooks, jobs scheduler   -->    classic: VMs in YOUR cloud account
  cluster manager                            serverless: in the DATABRICKS account
  Unity Catalog metastore service
                                          Data: your S3 / ADLS / GCS buckets

Control plane

This is the backend that Databricks operates. It holds the workspace web application, the notebook and job definitions, the cluster manager that starts and stops machines, and the metastore service for Unity Catalog. You do not run any of these.

Compute plane

This is where Spark and SQL actually run.

  • Classic compute: Databricks starts virtual machines inside your own cloud account (your VPC or VNet). You see the VMs in your cloud bill, plus Databricks charges.
  • Serverless compute: Databricks runs the compute in its own account, and you do not manage clusters at all. It starts in seconds, and you pay one combined price.

Data stays with you

Tables are files, usually Delta, in storage you own, such as S3, ADLS or GCS. Databricks does not hold your data. This is a major point in the lakehouse idea: the storage is open, so other engines can read the same files.

Databricks Runtime

The clusters run Databricks Runtime, which is Apache Spark plus Databricks's own optimizations (for example Photon and a faster Delta Lake implementation), and extra libraries for Python, ML and more. The runtime version is something you choose when you create compute, so jobs can be pinned to a tested version.

Why this matters in interviews

It explains security conversations (what leaves your network, what does not), networking questions (why classic compute needs VPC setup), and cost questions (VM cost plus DBU cost for classic, and a single price for serverless). Mention the data staying in your own storage first, since it is the point people care about.

PreviousNext