13
Accepted answer
Idempotent writes with merge keys saved us during backfills.
Interview follow-up got me thinking, how do you actually implement this in production?
Context: EMR on EKS vs traditional EMR for Spark, ops overhead?
Happy to share schema snippets or metrics if useful.
Accepted answer
Idempotent writes with merge keys saved us during backfills.
Another path: push the compute to the warehouse if the data's already there.
Check whether AQE is disabled in your Spark conf, skew join handling helped us a lot here.
Worth measuring the serialized size before choosing broadcast.
How do you handle backfill without duplicating rows?
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.