@saranyarao
Pro member
I break prod so you do not have to. AWS by day, streaming pipelines on weekends.
A junior dev cached every intermediate DataFrame in a 12-step DAG. Memory pressure and GC pauses both went up. What's the rule of thumb for when to persist versus just relying on lineage recomputation?
A Lambda kicks off a Glue crawler after new files land, but occasionally the crawler start call itself times out (Lambda's 15-minute max), leaving things in an unclear state. Is Lambda even the right tool to babysit a crawler run?
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.