Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Spot and preemptible instances for data jobs

Cloud · Cloud Architecture Basics

Spot and preemptible instances for data jobs

Mediumcloud-46
cloud-computespot-instancespreemptible-vmssparkcost-optimization

Question

When is it safe to use spot (preemptible) VMs for data processing?

Solution

Spot instances on AWS and preemptible virtual machines on Google Cloud provide 60 to 90 percent discounts compared to standard on-demand pricing, but cloud providers can reclaim them with only a few minutes of notice. They are safe to use for horizontally distributed, fault-tolerant batch workloads like Apache Spark executor nodes that can withstand individual worker terminations without crashing the entire job. However, cluster drivers and master nodes must remain on on-demand instances, intermediate state must be checkpointed, and spot nodes should never be used for latency-sensitive streaming or monolithic single-node jobs.

Safe workloads and cluster topology

Preemptible instances offer major cost savings if your system architecture is designed to handle node failures gracefully.

Cluster Master / Spark Driver  ---> On-Demand Instance  (Guaranteed Stability)
Worker Fleet / Spark Executors ---> Spot / Preemptible   (60-90% Discount, Retries)

Applying spot instances safely requires configuring heterogeneous cluster topologies:

  • Run the Spark driver and cluster master nodes on reliable on-demand instances. If a driver node gets reclaimed by the cloud provider, the entire Spark job fails immediately regardless of worker progress.
  • Deploy stateless Spark executors on spot instances. When a cloud provider reclaims a spot worker, Spark marks the executor lost, re-executes the missing partition tasks on remaining nodes, and continues processing without human intervention.
  • Enable automatic instance diversification across multiple instance types and availability zones in your autoscaling groups to minimize the chance that a capacity shortage reclaims all workers simultaneously.
  • Incorporate intermediate data checkpointing to object storage during long batch pipelines so a wave of preemption events does not force the pipeline to recompute the entire DAG from scratch.

When to avoid spot instances

Preemptible hardware introduces operational risks that make it unsuitable for specific workloads:

  • Never use spot instances for pipelines with strict delivery service level agreements, because repeated node preemptions introduce unpredictable task rescheduling delays.
  • Avoid spot instances for stateful streaming pipelines where losing nodes triggers continuous state rebalancing and stream consumer repartitioning.
  • Avoid spot compute for long single-node batch scripts that lack internal checkpointing and distributed retry mechanics.
PreviousNext