Prefer a staging table plus validation gate before promoting to prod tables.
Team is debating whether to rely on Redshift's automatic vacuum and analyze, or keep a manual scheduled job. Automatic sometimes runs at inconvenient times during business-hour query load.
Prefer a staging table plus validation gate before promoting to prod tables.
+1, saw identical behaviour after upgrading Spark 3.4 to 3.5.
ELI5: think of it like a phone book. If it's sorted by last name and you search by last name, that's fast. Search by first name instead and you're flipping through every page.
In simple terms: your job is asking for way more data than it actually needs, and that extra data has to travel across the network or between machines, which is slow. Ask for less, or ask smarter.
Auto vacuum delete and auto analyze are usually fine to leave on for most workloads, they're throttled to avoid heavy contention. Manual scheduling mainly helps when you know a specific low-traffic window and want a guaranteed full vacuum right after a big bulk load.
We replaced custom sensors with data contracts and row count checks.
+1, saw identical behaviour after upgrading Spark 3.4 to 3.5.
Small nit: the broadcast hint gets ignored once the table is over threshold, check the UI to confirm.
Careful with NULL in join keys, they'll drop rows in an inner join.
In simple terms: your job is asking for way more data than it actually needs, and that extra data has to travel across the network or between machines, which is slow. Ask for less, or ask smarter.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Short version for anyone skimming: this is a classic case of the tool doing exactly what you told it to, not what you meant. Double check the assumption, not the code.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.