Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. cache vs persist

PySpark · Caching & Reliability

cache vs persist

Hardpyspark-26
cachepersistmemory

Question

What is the difference between cache() and persist() in PySpark?

Solution

Both store a DataFrame/RDD so later actions reuse computed data instead of recomputing lineage.

cache()

Shortcut for persist() with default storage level:

  • DataFrames: typically MEMORY_AND_DISK (deserialized) in modern Spark SQL
  • Exact default has varied by API/version; say "cache is persist with the default level"

persist(level)

Lets you choose storage level explicitly.

from pyspark import StorageLevel

df.cache()
df.persist(StorageLevel.MEMORY_AND_DISK)
df.persist(StorageLevel.MEMORY_ONLY_SER)   # RDD-oriented example
df.unpersist()

Diagram

Action 1: compute + store partitions in memory/disk
Action 2: read from cache if still present
unpersist(): free storage

Gotcha

cache/persist are lazy markers. Materialization happens on the next action. Call an action after persist, or use df.count() / write, before relying on reuse.

Interview tip

Prefer caching only when the same DataFrame is reused multiple times; otherwise you waste memory.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext