Both tell Spark to reuse a computed DataFrame/RDD instead of recomputing its lineage every time an action needs it.
`cache()` is shorthand for persist with the default storage level (typically memory for DataFrames).
`persist(level)` lets you choose where data lives: memory only, memory+disk, disk only, serialized, etc.
df_clean = expensive_pipeline(df) df_clean.cache() # or persist(MEMORY_AND_DISK) df_clean.count() # materializes and stores df_clean.groupBy(...).agg(...) df_clean.filter(...).write... # without cache, lineage would recompute twice more
When to use
- Same DataFrame used by multiple downstream actions
- Iterative algorithms
When not to
- One-shot pipelines (wastes memory)
- Huge datasets that evict constantly (thrashing)
Always unpersist() when done in long-lived apps.
Interview tip: "cache = persist with default level. Use when reused; memory is not free."