Accepted answer
Document the grain decision, most BI bugs turn out to be grain bugs.
I'm the only person with data engineering skills on a 15-person startup, wearing analyst and DE hats at the same time. Every team wants their dashboard or pipeline yesterday. How do you triage fairly without becoming a full-time ticket queue?
Accepted answer
Document the grain decision, most BI bugs turn out to be grain bugs.
A visible, simple prioritization framework, even just impact versus effort on a shared doc, that stakeholders can actually see makes triage decisions feel less arbitrary and reduces the 'why is my thing not done yet' pressure, even when the answer is still 'not yet.'
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Consider DuckDB or Polars for this size before spinning up a cluster.
Idempotent writes with merge keys saved us during backfills.
The visible framework alone changed a lot of the friction. People stopped assuming their request was being ignored once they could see where it sat relative to others.
Any downside to this approach with incremental models?
Small nit: the broadcast hint gets ignored once the table is over threshold, check the UI to confirm.
Protect time explicitly for platform and infrastructure work, not just feature requests. As the only DE, if 100% of your time goes to the loudest requester, the foundational work that would make future requests faster never happens, and the treadmill only gets worse.
Idempotent writes with merge keys saved us during backfills.
Worth measuring the serialized size before choosing broadcast.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.