Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is idempotency?

Data platform · Pipelines

What is idempotency?

Mediumplatform-03
idempotencyretriesmergeexactly-once

Question

What does idempotency mean for data pipelines?

Solution

A pipeline step is idempotent if running it once or many times with the same input leaves the destination in the same correct state. Retries should not create duplicate facts or double-counted metrics.

Why you need it

Jobs fail halfway. Schedulers retry. Kafka consumers redeliver. Without idempotency, you get duplicated orders and inflated revenue.

Tiny bad example

-- NOT idempotent if run twice for the same day
insert into fct_orders
select * from staging_orders_2026_09_05;

Better patterns

  • Delete/replace a partition, then insert
  • Merge/upsert on a primary key
  • Use deterministic surrogate keys and unique constraints
  • Write to a dated partition path that overwrites
run 1: load day=2026-09-05 partition  -> 100 rows
run 2: overwrite same partition       -> still 100 rows (safe)

Interview tip: Define idempotency, show a non-idempotent insert, then name overwrite/merge as fixes.

PreviousNext