A backfill reprocesses historical data, usually after a bug fix, a new column, or a logic change. You rerun the pipeline for past dates so curated tables catch up.
bug found in tax logic on Sept 1 fix merges on Sept 10 backfill: re-run transforms for Sept 1..9
Safe backfill habits
1. Make the job date-partitioned / parameterized (run_date). 2. Make writes idempotent (replace partition or merge). 3. Bound the range; do not accidentally full-scan five years on day one. 4. Watch cost and cluster load; throttle parallelism. 5. Validate with row counts, checksums, or spot tests before flipping BI over.
Streaming note
Backfills may replay from Kafka retention, lake raw files, or source APIs. Retention limits can force you to keep a raw lake as the system of replay.
Interview tip: Lead with parameterized, idempotent, cost-aware replay. Do not say "just rerun prod."