At-least-once into Kafka with idempotent sink keys is pragmatic. Mention a dedup window or upsert keys in the serving layer and you're covered.
Interview prompt: 50k events/sec, 7-day raw retention, hourly aggregates for product dashboards, allow ad-hoc SQL on the last 24h.
How would you structure ingestion, storage layers, and serving without over-engineering it in an interview setting?
At-least-once into Kafka with idempotent sink keys is pragmatic. Mention a dedup window or upsert keys in the serving layer and you're covered.
For anyone who wants the two-sentence version: data arrived out of order, and the code assumed it wouldn't. Fix the assumption or sort the data, pick one.
Careful with NULL in join keys, they'll drop rows in an inner join.
Start with the execution plan, numbers beat guesses.
In our case the root cause was an implicit cast preventing pushdown.
ELI5 version: it's not broken, it's just slow because it's checking way more stuff than it needs to. Narrowing what it checks is almost always the fix.
Idempotent writes with merge keys saved us during backfills.
For anyone who wants the two-sentence version: data arrived out of order, and the code assumed it wouldn't. Fix the assumption or sort the data, pick one.
Any downside to this approach with incremental models?
Interviewer pushed on exactly-once vs at-least-once. What's a sane default answer there?
Any downside to this approach with incremental models?
Agree on the staging table swap. Atomic promote prevented partial reads.
Sign in to reply.
© 2026 Lakebench, operated by Hunnurji Rao. Bengaluru, Karnataka, India.
No cluster. No install. Just the tab.