In our case the root cause was an implicit cast preventing pushdown.
Migrated stacks last quarter and this regression appeared. Stack trace and plan attached:
Context: Dataflow streaming with BigQuery sink, duplicate rows on retry
Happy to share schema snippets or metrics if useful.
In our case the root cause was an implicit cast preventing pushdown.
Prefer a staging table plus validation gate before promoting to prod tables.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
In simple terms: your job is asking for way more data than it actually needs, and that extra data has to travel across the network or between machines, which is slow. Ask for less, or ask smarter.
Yes, the partial index was the win for us too.
We saw the same issue, fixing the partition filter dropped runtime 60%.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.