A checkpoint writes intermediate data to reliable storage and cuts the lineage DAG. After checkpointing, Spark no longer recomputes from the original sources for that Dataset; it reloads from the checkpoint directory if needed.
Cache vs checkpoint
cache/persist: - keeps data in memory/disk on executors - lineage retained (can recompute if lost) - not a durability boundary across app failures checkpoint: - writes to reliable storage (HDFS/S3/local checkpoint dir) - lineage truncated - used for long lineages / iterative algorithms / streaming state recovery patterns
Example (batch)
spark.sparkContext.setCheckpointDir("s3://bucket/spark-checkpoints/")
df2 = very_long_lineage_df.checkpoint() # eager checkpoint
# or df2 = very_long_lineage_df.localCheckpoint() # truncates but less reliableStreaming note
Structured Streaming uses checkpoint locations for offsets and state store recovery; that is related in spirit (fault tolerance) but configured on writeStream.option("checkpointLocation", ...).
Interview tip
Use checkpoint when lineage is deep enough that recomputation or planner overhead becomes painful, or when you need a recovery boundary. Prefer writing a curated intermediate table for most ETL.