Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. System Design

System design interview

Exactly-Once Processing in Streaming

MediumFree75 min read

Explain effectively-exactly-once processing: idempotent writes, at-least-once delivery, transactional APIs, and offset management.

system-designstreamingexactly-onceidempotency

The interview setup

Explain how to achieve effectively exactly-once processing in a distributed streaming pipeline. Discuss idempotent writes, at-least-once delivery, transactional APIs, and consumer offset management.

This is a concept interview as much as a design interview. Weak answers chant "Kafka transactions" and stop. Strong answers start with a double-charge story, define the three delivery guarantees, then show how source + processor + sink compose into an outcome that does not double-apply.

What the interviewer wants:

  • Precise definitions without hand-waving
  • A crash-between-side-effect-and-commit story
  • Idempotency keys and transactional boundaries
  • Honesty that "exactly-once" is often "effectively once"
  • A plan when the sink is a non-transactional external API

Background concepts from first principles

Why duplicates happen

Distributed systems fail in the middle. A consumer writes a row, then crashes before it commits the Kafka offset. On restart it reads the same message and writes again. The broker did its job with at-least-once delivery. Your sink was not prepared.

Duplicates also come from producer retries, rebalances, and replay.

The three guarantees

At-most-once: process then... actually, you may lose data if you commit offsets before finishing work, or drop on error. You try not to double-apply by accepting loss.

At-least-once: do not lose after durable accept; may double-apply on retries.

Exactly-once / effectively-once: duplicates do not change business outcomes. Either duplicates are prevented by transactions, or they are collapsed by idempotency, or both.

Offsets are commits of progress

In Kafka-style systems, the consumer offset is "how far I claim to have processed." If you commit early, you can lose. If you commit late, you can redo. The dance between side effects and offset commits is the whole game.

Idempotency

An operation is idempotent if doing it twice leaves the same result as doing it once. PUT /orders/123 with the same body is easier to make idempotent than POST /orders that always creates a new id. Design keys so retries collapse.

Transactions

A transaction makes a set of writes atomic. Kafka transactions can tie consumed offsets and produced sink records together. Flink checkpoints coordinate state and sinks with supported connectors. Transactions have a scope. Side effects outside the scope are not protected.

The expanded problem

Teach and design for a pipeline shaped like:

Source log (Kafka)
    -> stream processor (stateful or stateless)
    -> sink (Kafka / DB / warehouse / HTTP API)

Cover:

  • Definitions of the three guarantees
  • Why naive consumers double-write
  • Idempotent sinks
  • Transactional produce/consume patterns
  • Offset management
  • Warehouse merge semantics
  • External API without transactions

What good looks like

  • Start with a story, then definitions
  • Draw the crash window between sink write and offset commit
  • Separate tool features from end-to-end outcomes
  • Give a concrete answer for DB sinks and for HTTP sinks
  • Mention non-determinism as an idempotency killer

Clarifying questions

  • What is the sink: Kafka topic, relational DB, object-store table, or external SaaS API?
  • Are duplicate outputs visible to end users or only to internal tables?
  • Is processing deterministic given an input event?
  • Do we need strict EOS or is idempotent at-least-once enough?
  • Multi-destination sinks in one job?

Out of scope

  • Formal consensus protocol proofs
  • Implementing a database kernel
  • Claiming cosmic EOS across email, SMS, and Salesforce without caveats

How to use clarifying questions

Ask a few questions that change architecture: SLA definition, source of truth, retention, and tolerance for approximates. If told to decide, state assumptions on the board and proceed.

What good sounds like

Narrate trade-offs while drawing. Name the waiting room for bursts, the transform, the serving surface, and the replay path. Explicitly list out-of-scope items.

Before you draw

Find entry, buffer, transform, serve, and replay. If any is missing, your operational story has a hole.

How to use clarifying questions

Ask a few questions that change architecture: SLA definition, source of truth, retention, and tolerance for approximates. If told to decide, state assumptions on the board and proceed.

What good sounds like

Narrate trade-offs while drawing. Name the waiting room for bursts, the transform, the serving surface, and the replay path. Explicitly list out-of-scope items.

Before you draw

Find entry, buffer, transform, serve, and replay. If any is missing, your operational story has a hole.

How to use clarifying questions

Ask a few questions that change architecture: SLA definition, source of truth, retention, and tolerance for approximates. If told to decide, state assumptions on the board and proceed.

What good sounds like

Narrate trade-offs while drawing. Name the waiting room for bursts, the transform, the serving surface, and the replay path. Explicitly list out-of-scope items.

Before you draw

Find entry, buffer, transform, serve, and replay. If any is missing, your operational story has a hole.

How to use clarifying questions

Ask a few questions that change architecture: SLA definition, source of truth, retention, and tolerance for approximates. If told to decide, state assumptions on the board and proceed.

What good sounds like

Narrate trade-offs while drawing. Name the waiting room for bursts, the transform, the serving surface, and the replay path. Explicitly list out-of-scope items.

Before you draw

Find entry, buffer, transform, serve, and replay. If any is missing, your operational story has a hole.

Why this question shows up so often

Data engineers live in the overlap of unreliable networks and stateful business systems. Hiring teams use exactly-once questions to see whether you will accidentally double-charge customers, double-count metrics, or claim impossible guarantees across SaaS boundaries. Treat it as a reliability design problem, not trivia.