Accepted answer
Cache when something is reused across branches or in iterative logic. For single-use transforms, just let Catalyst recompute.
A junior dev cached every intermediate DataFrame in a 12-step DAG. Memory pressure and GC pauses both went up.
What's the rule of thumb for when to persist versus just relying on lineage recomputation?
Accepted answer
Cache when something is reused across branches or in iterative logic. For single-use transforms, just let Catalyst recompute.
Check the Storage tab: MEMORY_AND_DISK with spill usually means you cached too much. Unpersist aggressively.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
For long lineage cuts on wide DAGs, write a checkpoint to reliable storage instead of caching.
ELI5 version: it's not broken, it's just slow because it's checking way more stuff than it needs to. Narrowing what it checks is almost always the fix.
Simple way to think about it: caching is a bet that you'll ask the same question again soon. If you don't, you're just paying rent on memory for nothing.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Event-driven beats cron once landing time gets unpredictable.
Document the grain decision, most BI bugs turn out to be grain bugs.
ELI5: think of it like a phone book. If it's sorted by last name and you search by last name, that's fast. Search by first name instead and you're flipping through every page.
Consider DuckDB or Polars for this size before spinning up a cluster.
Agree on the staging table swap. Atomic promote prevented partial reads.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.