Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is a checkpoint?

PySpark · Caching & Reliability

What is a checkpoint?

Hardpyspark-28
checkpointlineagefault-tolerance

Question

What is a checkpoint in Spark, and how does it differ from cache?

Solution

A checkpoint writes intermediate data to reliable storage and cuts the lineage DAG. After checkpointing, Spark no longer recomputes from the original sources for that Dataset; it reloads from the checkpoint directory if needed.

Cache vs checkpoint

cache/persist:
  - keeps data in memory/disk on executors
  - lineage retained (can recompute if lost)
  - not a durability boundary across app failures

checkpoint:
  - writes to reliable storage (HDFS/S3/local checkpoint dir)
  - lineage truncated
  - used for long lineages / iterative algorithms / streaming state recovery patterns

Example (batch)

spark.sparkContext.setCheckpointDir("s3://bucket/spark-checkpoints/")

df2 = very_long_lineage_df.checkpoint()  # eager checkpoint
# or df2 = very_long_lineage_df.localCheckpoint()  # truncates but less reliable

Streaming note

Structured Streaming uses checkpoint locations for offsets and state store recovery; that is related in spirit (fault tolerance) but configured on writeStream.option("checkpointLocation", ...).

Interview tip

Use checkpoint when lineage is deep enough that recomputation or planner overhead becomes painful, or when you need a recovery boundary. Prefer writing a curated intermediate table for most ETL.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext