Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
I have read the docs but real-world tradeoffs are unclear. What would you optimize first?
Context: Guaranteeing order per user_id across partitions
Happy to share schema snippets or metrics if useful.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Idempotent writes with merge keys saved us during backfills.
Any downside to this approach with incremental models?
Note that merge on Delta still needs unique keys defined correctly.
In our case the root cause was an implicit cast preventing pushdown.
Any downside to this approach with incremental models?
Document the grain decision, most BI bugs turn out to be grain bugs.
Careful with NULL in join keys, they'll drop rows in an inner join.
Prefer a staging table plus validation gate before promoting to prod tables.
ELI5 version: it's not broken, it's just slow because it's checking way more stuff than it needs to. Narrowing what it checks is almost always the fix.
Short version for anyone skimming: this is a classic case of the tool doing exactly what you told it to, not what you meant. Double check the assumption, not the code.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.