Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Cache vs persist

Batch & Streaming · Batch Processing

Cache vs persist

Mediumstream-18
cachepersistmemorylineage

Question

What is the difference between cache and persist in Spark?

Solution

Both tell Spark to reuse a computed DataFrame/RDD instead of recomputing its lineage every time an action needs it.

`cache()` is shorthand for persist with the default storage level (typically memory for DataFrames).

`persist(level)` lets you choose where data lives: memory only, memory+disk, disk only, serialized, etc.

df_clean = expensive_pipeline(df)
df_clean.cache()          # or persist(MEMORY_AND_DISK)

df_clean.count()          # materializes and stores
df_clean.groupBy(...).agg(...)
df_clean.filter(...).write...
# without cache, lineage would recompute twice more

When to use

  • Same DataFrame used by multiple downstream actions
  • Iterative algorithms

When not to

  • One-shot pipelines (wastes memory)
  • Huge datasets that evict constantly (thrashing)

Always unpersist() when done in long-lived apps.

Interview tip: "cache = persist with default level. Use when reused; memory is not free."

PreviousNext