Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is a backfill?

Pipelines & scenarios · Core Pipeline Concepts

What is a backfill?

Mediumpipe-15
backfillreprocessingpartitionsopsidempotency

Question

What is a data backfill, and how do you run one safely?

Solution

A backfill reprocesses historical data after a bug fix, new column, or logic change so curated tables catch up for past dates.

Bug in tax logic from Sept 1
Fix merges on Sept 10

Backfill: re-run transforms for Sept 1 .. Sept 9
(parameterized by run_date, idempotent writes)

Safe backfill habits

1. Parameterize the job (run_date / interval). 2. Ensure writes are idempotent (replace partition or merge). 3. Bound the range; do not full-scan five years on day one. 4. Throttle parallelism to protect the warehouse/cluster cost. 5. Validate row counts / checksums before flipping BI.

Backfill vs catchup

  • Backfill: intentional historical replay you control.
  • Catchup: orchestrator automatically filling missed scheduled intervals.

Streaming note

Replay from Kafka retention, raw lake files, or source APIs. If retention is short, the raw lake is your system of replay.

Interview tip: Emphasize parameterized, idempotent, cost-aware replay, not "just rerun prod for all time."

PreviousNext