Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is exactly-once processing?

PySpark · UDFs & Streaming

What is exactly-once processing?

Hardpyspark-35
exactly-oncestreamingcheckpointfault-tolerance

Question

What does exactly-once processing mean in Spark Structured Streaming?

Solution

Exactly-once means each input record affects the result once, even if the job retries after failure. No silent duplicates, no silent drops (for the guaranteed sink/source combinations).

How Spark approaches it

checkpoint (offsets + state)
   + idempotent / transactional sink
   = end-to-end exactly-once (when supported)

1. Checkpoint location stores offsets and state so replay starts at the right place. 2. On retry, Spark recomputes a micro-batch using those offsets. 3. The sink must not double-apply side effects. File sinks use commit protocols; Kafka / Delta / Iceberg have specific guarantees; arbitrary foreach may be at-least-once unless you make writes idempotent.

Modes in practice

  • At-least-once: easier; duplicates possible on retry
  • Exactly-once end-to-end: requires replayable source + checkpoint + supporting sink
  • Effectively exactly-once: idempotent keys / merge into a table (dedupe by id)
(
    df.writeStream
    .format("delta")  # example of a transactional sink ecosystem
    .option("checkpointLocation", "/chk/orders")
    .outputMode("append")
    .start("/lake/silver/orders")
)

Interview tip

Be precise: Spark can provide exactly-once within its engine semantics for supported sinks; external systems need idempotency or transactions to make the business effect exactly-once.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext