A pipeline step is idempotent if running it once or many times with the same input leaves the destination in the same correct state. Retries should not create duplicate facts or double-counted metrics.
Why you need it
Jobs fail halfway. Schedulers retry. Kafka consumers redeliver. Without idempotency, you get duplicated orders and inflated revenue.
Tiny bad example
-- NOT idempotent if run twice for the same day insert into fct_orders select * from staging_orders_2026_09_05;
Better patterns
- Delete/replace a partition, then insert
- Merge/upsert on a primary key
- Use deterministic surrogate keys and unique constraints
- Write to a dated partition path that overwrites
run 1: load day=2026-09-05 partition -> 100 rows run 2: overwrite same partition -> still 100 rows (safe)
Interview tip: Define idempotency, show a non-idempotent insert, then name overwrite/merge as fixes.