Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Design a safely retryable pipeline

Pipelines & scenarios · Design & Scenarios

Design a safely retryable pipeline

Hardpipe-21
retriesidempotencyatomic publishdesignreliability

Question

How would you design a pipeline that is safe to retry?

Solution

A safely retryable pipeline can be re-run after failure without duplicating or corrupting outputs. Idempotency is the core design goal.

Pattern: stage → validate → atomic publish

  1) write to temp / staging path or staging table
  2) run DQ checks
  3) swap / MERGE overwrite target partition
     (or MERGE on keys)

Design rules

1. Partition or key your writes (dt=... or business primary key). 2. Prefer overwrite/merge over blind append. 3. Commit offsets / watermarks only after success. 4. Make side effects safe (emails, webhooks: dedupe tokens). 5. Keep raw inputs durable so you can replay.

Bad vs good

BAD:  INSERT into fct; commit Kafka offset; crash; retry inserts again
GOOD: write staging → merge fct → then commit offset

Orchestrator settings

Retries with backoff for transient errors; fail and alert for poison data after N tries; quarantine bad files instead of blocking forever.

Interview tip: Say "retryable = idempotent outputs + late watermark/offset commit + atomic publish." That sentence is the answer.

PreviousNext