Worth measuring the serialized size before choosing broadcast.
Building a Flink job reading from Kafka, writing aggregates back to Kafka. Team is debating whether enabling Kafka transactions (EXACTLY_ONCE sink semantics) is worth the throughput cost versus at-least-once with idempotent downstream consumers.
Worth measuring the serialized size before choosing broadcast.
Our aggregate is a running count, which isn't naturally idempotent. Sounds like we do need it here.
We replaced custom sensors with data contracts and row count checks.
If downstream consumers can dedupe on a natural key, at-least-once end to end is usually simpler and faster than paying the transaction coordination overhead for true exactly-once semantics.
We saw the same issue, fixing the partition filter dropped runtime 60%.
Worth measuring the serialized size before choosing broadcast.
Start with the execution plan, numbers beat guesses.
+1, saw identical behaviour after upgrading Spark 3.4 to 3.5.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.