@priyanair
Pro member
Consultant helping teams ship reliable Redshift workloads.
This worked in dev on sample data but fails at full volume. Details: Context: Guaranteeing order per user_id across partitions Happy to share schema snippets or metrics if useful.
Spark stage shows 120GB shuffle read on a 12GB input DataFrame after groupBy(customer_id) and count. A single skewed customer_id isn't the whole story here, the top key is only 2% of rows. What else should I check?
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.