Both store a DataFrame/RDD so later actions reuse computed data instead of recomputing lineage.
cache()
Shortcut for persist() with default storage level:
- DataFrames: typically MEMORY_AND_DISK (deserialized) in modern Spark SQL
- Exact default has varied by API/version; say "cache is persist with the default level"
persist(level)
Lets you choose storage level explicitly.
from pyspark import StorageLevel df.cache() df.persist(StorageLevel.MEMORY_AND_DISK) df.persist(StorageLevel.MEMORY_ONLY_SER) # RDD-oriented example df.unpersist()
Diagram
Action 1: compute + store partitions in memory/disk Action 2: read from cache if still present unpersist(): free storage
Gotcha
cache/persist are lazy markers. Materialization happens on the next action. Call an action after persist, or use df.count() / write, before relying on reuse.
Interview tip
Prefer caching only when the same DataFrame is reused multiple times; otherwise you waste memory.