Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
This worked in dev on sample data but fails at full volume. Details:
Context: SQL window function question, top 3 orders per customer
Happy to share schema snippets or metrics if useful.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
What batch size or interval worked for you at similar scale?
Careful with NULL in join keys, they'll drop rows in an inner join.
Sometimes the real fix is a product change so you stop needing that join at all.
Worth measuring the serialized size before choosing broadcast.
Event-driven beats cron once landing time gets unpredictable.
Sometimes the real fix is a product change so you stop needing that join at all.
We replaced custom sensors with data contracts and row count checks.
How do you handle backfill without duplicating rows?
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.