A job is safe to rerun when running it twice gives the same result as running it once. You get there by making outputs predictable, writing them atomically, and recording what has been done.
Deterministic output location
Derive the output path from the logical run date, not from the current time or a random id. The run for 2025-03-01 always writes to orders/run_date=2025-03-01/. A rerun overwrites that same location instead of adding a second copy somewhere else.
Write to a temporary name, then rename
import os, tempfile
def write_atomic(df, final_path):
tmp = final_path + ".tmp"
df.to_parquet(tmp)
os.replace(tmp, final_path) # atomic on the same filesystemReaders never see a half-written file, and a crash during writing leaves only the temporary file, which the next run cleans up. On object storage, rename is not truly atomic, so write to a staging prefix and commit by writing a small success marker file or updating a manifest, and have readers trust the marker.
Upsert by key
When loading into a table, use MERGE or INSERT ... ON CONFLICT DO UPDATE keyed on a business key, not a plain INSERT. Reloading the same rows then changes nothing.
Track what is processed
For file-driven jobs, keep a manifest: a table or file recording each input file (name, size, checksum) and its status. Before processing a file, check the manifest. After a successful load, record it, ideally in the same transaction as the load. A restart then skips what is done and resumes what is not.
No side effects before success
Do the risky and visible things last. Do not send an email, move the source file, or update a "last run" marker until the data has been written and verified. If you must call an external system, use an idempotency key so a repeat is ignored.
Same input, same output
Avoid anything that makes output change from run to run: the current time inside the transformation (pass it in as a parameter), random numbers without a seed, and unordered results used for logic. Then a backfill of last month gives the same numbers as the original run.
Quick test
Run the job twice in a row on the same input and compare the outputs. If they differ, it is not idempotent yet.