Accepted answer
Document the grain decision, most BI bugs turn out to be grain bugs.
An ETL job built on top of an ORM for convenience turned out to be issuing one query per row for a related lookup, invisible until volume grew. 10k rows went from a few seconds to 20+ minutes.
for order in orders:
customer = Customer.query.get(order.customer_id)How do you catch this class of bug before it hits production volume?
Accepted answer
Document the grain decision, most BI bugs turn out to be grain bugs.
Agree on the staging table swap. Atomic promote prevented partial reads.
Use eager loading, joinedload or select_related depending on the ORM, explicitly for any relation accessed in a loop, and add a query-count assertion in tests. Most ORMs expose a query counter you can assert against to catch N+1 regressions in CI, not just in code review.
What batch size or interval worked for you at similar scale?
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.