Accepted answer
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Added an index expecting a speedup on a filtered query, but EXPLAIN ANALYZE shows it actually runs slower than the original sequential scan.
CREATE INDEX idx_orders_status ON orders(status);status has only 4 distinct values on a 20M row table.
Accepted answer
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Prefer a staging table plus validation gate before promoting to prod tables.
Idempotent writes with merge keys saved us during backfills.
Low-cardinality columns are the classic case where an index makes things worse. The planner, correctly, still often prefers a seq scan since a huge fraction of rows match any given status value, and the extra index lookup overhead isn't worth it.
Another path: push the compute to the warehouse if the data's already there.
How do you handle backfill without duplicating rows?
This matches our runbook for skewed keys.
Sometimes the real fix is a product change so you stop needing that join at all.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.