Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Separation of storage and compute

Snowflake, BigQuery & Databricks · Choosing and Comparing

Separation of storage and compute

Easywarehouses-54
storage-compute-separationmppscalingworkload-isolation

Question

Why does separating storage and compute matter in modern warehouses?

Solution

Separating storage from compute means data lives in shared cloud object storage, and the machines that process it are independent. You can scale, start and stop each side separately, and many compute clusters can use one copy of the data.

The old way

In classic MPP warehouses (Teradata, Netezza, Redshift DC2 nodes, on-premises Hadoop), each node held a slice of the data on its local disks and also ran the queries. That tied the two together:

  • To add storage you had to add nodes, paying for compute you did not need.
  • To add compute for a busy hour, you had to add nodes and redistribute data, which was slow.
  • All workloads shared the same nodes, so a heavy load job slowed dashboards.
  • You paid for the cluster all day, even when nothing ran.

The modern way

Object storage (S3 / GCS / ADLS)   <-- one copy of data
     ^          ^          ^
 ETL cluster  BI cluster  data-science cluster   <-- independent compute

What it gives you

  • Independent scaling: grow storage cheaply as data grows, and grow compute only for heavy periods.
  • Many clusters on the same data with no copies, so each team can have its own compute.
  • Pay for compute only while it runs. A cluster suspended overnight costs nothing, and storage stays cheap.
  • Workload isolation: an ETL job and a dashboard do not compete for CPU.
  • Faster recovery and resizing, since moving data is not required when you change cluster size.

The costs

Reading from object storage has more latency than local disks, so engines add caches (such as Snowflake's local disk cache) to hide it. Egress between regions or clouds can cost money. And because it is easy to start compute, spending can grow without anyone noticing, which makes monitoring and limits important.

Mention in interviews

Teradata or older Redshift as the "before", Snowflake, BigQuery and Databricks as the "after", and one example of isolation. Also say that it is the idea behind lakehouse formats, where open files in object storage are read by many engines.

PreviousNext