@zoyajoshi
Free member
Data engineer at a a SaaS analytics team. Mostly Iceberg and streaming pipelines.
A running total using Window.partitionBy("customer_id").orderBy("date") worked fine in dev on 1M rows, then fell over on 500M rows in prod with massive shuffle spill. Anything to do besides get a bigger cluster?
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.