Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. When should you cache?

PySpark · Caching & Reliability

When should you cache?

Hardpyspark-27
cacheperformancebest-practices

Question

When should you cache a DataFrame in Spark, and when should you avoid it?

Solution

Cache when

  • The DataFrame is reused by multiple downstream actions (several writes, ML feature reuse, iterative algorithms)
  • Recomputation is expensive (heavy join + aggregate) and the working set fits reasonably in memory/disk
  • Interactive exploration repeatedly queries the same cleaned subset
base = spark.read.parquet("/lake/silver/orders").filter("dt >= '2024-01-01'")
feat = expensive_joins(base).cache()
feat.count()  # materialize

feat.groupBy("region").count().show()
feat.write.mode("overwrite").parquet("/tmp/feat")
feat.unpersist()

Avoid caching when

  • Single-pass ETL: read -> transform -> write once
  • Dataset is huge relative to executor memory (constant eviction/spill)
  • You only need durability: write a checkpoint table/Parquet instead of memory cache
  • Caching would hide partition issues you should fix at the source

Diagram

Good:   heavy_df -> cache -> action A
                           -> action B

Bad:    df.cache(); df.write...   # single use, wasted memory

Interview tip

"Cache is a reuse optimization, not a correctness tool. For lineage truncation and fault isolation, use checkpoint or materialize to storage."

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext