Exactly-once means each input record affects the result once, even if the job retries after failure. No silent duplicates, no silent drops (for the guaranteed sink/source combinations).
How Spark approaches it
checkpoint (offsets + state) + idempotent / transactional sink = end-to-end exactly-once (when supported)
1. Checkpoint location stores offsets and state so replay starts at the right place. 2. On retry, Spark recomputes a micro-batch using those offsets. 3. The sink must not double-apply side effects. File sinks use commit protocols; Kafka / Delta / Iceberg have specific guarantees; arbitrary foreach may be at-least-once unless you make writes idempotent.
Modes in practice
- At-least-once: easier; duplicates possible on retry
- Exactly-once end-to-end: requires replayable source + checkpoint + supporting sink
- Effectively exactly-once: idempotent keys / merge into a table (dedupe by id)
(
df.writeStream
.format("delta") # example of a transactional sink ecosystem
.option("checkpointLocation", "/chk/orders")
.outputMode("append")
.start("/lake/silver/orders")
)Interview tip
Be precise: Spark can provide exactly-once within its engine semantics for supported sinks; external systems need idempotency or transactions to make the business effect exactly-once.