Agree on the staging table swap. Atomic promote prevented partial reads.
A few models share a nearly-identical 10-line Jinja block for a common date-bucketing pattern. Feels like it should be a macro, but the models each have slightly different edge cases. When do you extract, and when do you just accept duplication?
Agree on the staging table swap. Atomic promote prevented partial reads.
ELI5: think of it like a phone book. If it's sorted by last name and you search by last name, that's fast. Search by first name instead and you're flipping through every page.
Our dimension is slowly changing, does that change the join order?
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
In our case the root cause was an implicit cast preventing pushdown.
Consider DuckDB or Polars for this size before spinning up a cluster.
Prefer a staging table plus validation gate before promoting to prod tables.
Careful with NULL in join keys, they'll drop rows in an inner join.
Start with the execution plan, numbers beat guesses.
This matches our runbook for skewed keys.
Start with the execution plan, numbers beat guesses.
Yes, the partial index was the win for us too.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.