Accepted answer
Worth measuring the serialized size before choosing broadcast.
This worked in dev on sample data but fails at full volume. Details:
Context: CDC from OLTP to lakehouse, Debezium vs DMS
Happy to share schema snippets or metrics if useful.
Accepted answer
Worth measuring the serialized size before choosing broadcast.
In our case the root cause was an implicit cast preventing pushdown.
Another path: push the compute to the warehouse if the data's already there.
Idempotent writes with merge keys saved us during backfills.
+1, saw identical behaviour after upgrading Spark 3.4 to 3.5.
Any downside to this approach with incremental models?
Worth measuring the serialized size before choosing broadcast.
Our dimension is slowly changing, does that change the join order?
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.