Fault tolerance means the pipeline can survive machine crashes, transient network errors, and partial job failures without corrupting data or losing committed work permanently.
Job writing day=2026-09-05 task 1 OK task 2 crash mid-write task 3 OK Fault-tolerant design: - do not publish a half-written partition as "success" - retry failed tasks - overwrite or stage-then-swap so readers see all-or-nothing
Building blocks
- Retries with backoff for transient errors
- Checkpoints (Spark streaming, Flink state)
- Idempotent writes so retries are safe
- Transactional / atomic publishes (temp path + rename, Delta/Iceberg commits)
- Replayable sources (raw lake, Kafka retention)
Not the same as "never fails"
Fault-tolerant systems assume failure is normal and recover cleanly.
Interview tip: Tie fault tolerance to retries + checkpoints + idempotent/atomic writes. "We retry" alone is incomplete if retries double-write.