Document the grain decision, most BI bugs turn out to be grain bugs.
Composer environment started at ~40 DAGs with fast parsing. Now at ~300 DAGs, the scheduler's DAG parse loop takes minutes, and task scheduling latency has visibly degraded.
Document the grain decision, most BI bugs turn out to be grain bugs.
Event-driven beats cron once landing time gets unpredictable.
ELI5: think of it like a phone book. If it's sorted by last name and you search by last name, that's fast. Search by first name instead and you're flipping through every page.
Prefer a staging table plus validation gate before promoting to prod tables.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Another path: push the compute to the warehouse if the data's already there.
Careful with NULL in join keys, they'll drop rows in an inner join.
Check for expensive top-level code in your DAG files, API calls, DB queries, heavy imports executed at parse time rather than inside operators. That's the most common cause of parse-time blowup as DAG count scales, and it compounds because every DAG file re-executes on every scheduler loop.
ELI5: think of it like a phone book. If it's sorted by last name and you search by last name, that's fast. Search by first name instead and you're flipping through every page.
Small nit: the broadcast hint gets ignored once the table is over threshold, check the UI to confirm.
This matches our runbook for skewed keys.
Note that merge on Delta still needs unique keys defined correctly.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.