Cache when
- The DataFrame is reused by multiple downstream actions (several writes, ML feature reuse, iterative algorithms)
- Recomputation is expensive (heavy join + aggregate) and the working set fits reasonably in memory/disk
- Interactive exploration repeatedly queries the same cleaned subset
base = spark.read.parquet("/lake/silver/orders").filter("dt >= '2024-01-01'")
feat = expensive_joins(base).cache()
feat.count() # materialize
feat.groupBy("region").count().show()
feat.write.mode("overwrite").parquet("/tmp/feat")
feat.unpersist()Avoid caching when
- Single-pass ETL: read -> transform -> write once
- Dataset is huge relative to executor memory (constant eviction/spill)
- You only need durability: write a checkpoint table/Parquet instead of memory cache
- Caching would hide partition issues you should fix at the source
Diagram
Good: heavy_df -> cache -> action A
-> action B
Bad: df.cache(); df.write... # single use, wasted memoryInterview tip
"Cache is a reuse optimization, not a correctness tool. For lineage truncation and fault isolation, use checkpoint or materialize to storage."