A step is idempotent if running it once or many times with the same inputs leaves the destination in the same correct state. Retries must not double-count facts.
Why it matters
Jobs fail halfway. Orchestrators retry. Kafka redelivers. Without idempotency you get duplicate revenue and broken dashboards.
BAD (append-only, not idempotent): run 1 for day=2026-09-05 -> 100 rows retry same day -> 200 rows (wrong) GOOD (overwrite partition): run 1: replace day=2026-09-05 -> 100 rows retry: replace day=2026-09-05 -> still 100 rows
Practical patterns
- Delete/replace a partition, then insert
MERGE/ upsert on a natural key- Deterministic primary keys + unique constraints
- Write to a dated path that overwrites (
s3://.../dt=2026-09-05/)
Tiny SQL smell
-- NOT safe to retry blindly INSERT INTO fct_orders SELECT * FROM staging_orders_2026_09_05;
Interview tip: Define the word, show a non-idempotent insert, then name overwrite/merge as the fix. Interviewers listen for this habit.