Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Optimize cloud costs (5 strategies)

Cloud · General Cloud

Optimize cloud costs (5 strategies)

Mediumcloud-20
costfinopsoptimizationstoragecompute

Question

Name five practical strategies to optimize cloud data platform costs.

Solution

Cloud bills explode from idle clusters, full-table scans, hot storage of cold data, and chatty APIs. Here are five high-value strategies interviewers like:

1. Right-size and shut down idle compute

Stop/terminate EMR/Databricks clusters when jobs end. Prefer ephemeral jobs (Glue, job clusters) over always-on interactive clusters for production ETL. Use autoscaling.

2. Lifecycle object storage

Move old S3/GCS/ADLS data to IA / Nearline / Cool / Archive classes. Delete temp scratch prefixes. Partition so you can expire by date.

3. Reduce query scan bytes

Partition and cluster warehouse tables. Select only needed columns (columnar formats help). Avoid SELECT * on huge fact tables. Cache or materialize heavy marts.

4. Pick serverless vs provisioned by utilization

Spiky workloads → serverless / on-demand. Steady high utilization → reserved / provisioned capacity (savings plans, slot commitments, reserved instances).

5. Observe and allocate

Turn on cost tags / labels per team/env. Budgets + alerts. Review top jobs weekly (Spark stage metrics, BigQuery bytes billed, warehouse credit burn).

Before: always-on cluster + Standard storage forever + SELECT *
After:  job clusters + lifecycle rules + partitioned queries + tags/alerts

Interview tip: List five concrete levers (idle compute, storage tiering, scan reduction, pricing model fit, cost observability). One example each beats a vague "be careful."

PreviousNext