Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Warehouse costs spiked after a deploy

Pipelines & scenarios · Production Scenarios

Warehouse costs spiked after a deploy

Hardpipelines-51
scenariocostdeploydbtquery-history

Question

After a dbt or Spark deploy, warehouse cost jumped 4x. How do you find the cause?

Solution

A 4x cost jump right after a deploy is a strong clue, because it narrows the search to what changed. Compare before and after using query history, and find the query or model that account for the difference.

Step 1: compare the periods

In the warehouse's query history (Snowflake QUERY_HISTORY, BigQuery INFORMATION_SCHEMA.JOBS, Databricks system tables), group by day and by model, job or query tag. Find which jobs grew in credits, bytes scanned or slot time, and by how much. Usually one or two items explain most of the jump.

Step 2: usual causes after a deploy

  • An incremental model that now rebuilds everything (a config change, a changed unique key, or a --full-refresh flag left in the schedule). Look at the compile output or run logs for "full refresh".
  • A removed or broken partition filter, so a query scans the entire table instead of the last few days. Check bytes scanned per run before and after.
  • A new join that multiplies rows (a missing condition or a non-unique key), which shows as rows out far above rows in.
  • A model changed to a table materialization from view or incremental, or a view layered on a view that is now being queried by a dashboard.
  • Schedule changes: a job running every 5 minutes instead of hourly, or several overlapping runs.
  • A new dependency that makes models run in a different order, repeating work.
  • A larger warehouse size set in the deploy config.

Step 3: act

Roll back or hotfix the cause, since cost is running every hour. Then rerun and confirm the daily spend falls to the earlier level. If there was a full rebuild of large tables, note that part of the spike is a one-off.

Step 4: stop it happening again

  • Cost checks in CI: for dbt, compare a dry-run estimate or the bytes scanned of changed models in a pull request, and flag large increases.
  • Guardrails: a maximum bytes billed per query (BigQuery), resource monitors (Snowflake), cluster policies (Databricks).
  • Alerts on daily spend versus the 7-day average, per team or tag.
  • Review of materialization and schedule changes in pull requests.
  • Tag queries by job, so the next investigation takes minutes.

Present it as a method (compare before and after, rank by cost, find the change), and give two or three concrete examples of causes.

PreviousNext