7
Accepted answer
Start with the execution plan, numbers beat guesses.
Team is split on approach A vs B. Looking for experiences from similar scale.
Context: Dataflow streaming with BigQuery sink, duplicate rows on retry
Happy to share schema snippets or metrics if useful.
Accepted answer
Start with the execution plan, numbers beat guesses.
Prefer a staging table plus validation gate before promoting to prod tables.
Idempotent writes with merge keys saved us during backfills.
We saw the same issue, fixing the partition filter dropped runtime 60%.
Prefer a staging table plus validation gate before promoting to prod tables.
+1, saw identical behaviour after upgrading Spark 3.4 to 3.5.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.