Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. System Design

System design interview

Backfill Strategy for Historical Data

MediumPro55 min read

Reprocess 2 years after a bug fix using atomic swaps, incremental date ranges, without disrupting live consumers.

system-designbackfilldata-opslakehouse

The interview setup

Design a backfill to reprocess 2 years of history after a bug fix. Discuss atomic swap, incremental backfill by date, and avoiding disruption to live consumers.

Backfill is a publication problem. Wrong numbers already escaped. You must rebuild without torn reads and without fighting today's writer.

Background concepts from first principles

Why in-place overwrite is dangerous

BI tools query while you rewrite. Readers can see half-old half-new files. Or writers for today collide with historical rebuilds on the same partition.

Shadow then swap

Write corrected data to a shadow location or shadow table. Validate. Atomically replace partition or swap table metadata.

Incremental by date

Loop date partitions. Parallelize carefully. Isolate live ownership: live pipeline owns date >= cutover; backfill owns history.

Validation

Compare row counts, aggregates, and sampled keys against expectations before publish. Backfill without validation is how you ship a second bug faster.

The expanded problem

  • Rebuild affected tables for ~2 years
  • Publish atomically per partition or via table swap
  • Keep incremental live pipeline running
  • Validate before cutover
  • Bound blast radius and estimate cost/time
For date in history:
  transform --> shadow partition
validate samples / aggregates
atomic replace partition (or swap table)
live pipeline continues for today+

What good looks like

  • Bug impact scoping first
  • Shadow writes
  • Date loop with isolation from live
  • Validation gates
  • Consumer communication

Clarifying questions

  • Which downstreams read which partitions?
  • Bug in transform logic or source data?
  • Can consumers pause or read dual versions?
  • How far does the dependency graph extend (marts on marts)?

Out of scope

  • Fixing the original business logic bug itself in detail
  • Rebuilding the entire company warehouse unrelated tables

Clarifying questions strategy

Ask what changes grain, SLA, retention, or cost. State assumptions when answers are vague.

What good looks like

Grain or policy first, boxes second. Atomic publish. Safe reruns. Clear out-of-scope.

Scope control

Protect the critical path. Park optional marts after the SLA landing succeeds.

Clarifying questions strategy

Ask what changes grain, SLA, retention, or cost. State assumptions when answers are vague.

What good looks like

Grain or policy first, boxes second. Atomic publish. Safe reruns. Clear out-of-scope.

Scope control

Protect the critical path. Park optional marts after the SLA landing succeeds.

Clarifying questions strategy

Ask what changes grain, SLA, retention, or cost. State assumptions when answers are vague.

What good looks like

Grain or policy first, boxes second. Atomic publish. Safe reruns. Clear out-of-scope.

Scope control

Protect the critical path. Park optional marts after the SLA landing succeeds.

Clarifying questions strategy

Ask what changes grain, SLA, retention, or cost. State assumptions when answers are vague.

What good looks like

Grain or policy first, boxes second. Atomic publish. Safe reruns. Clear out-of-scope.

Scope control

Protect the critical path. Park optional marts after the SLA landing succeeds.

Clarifying questions strategy

Ask what changes grain, SLA, retention, or cost. State assumptions when answers are vague.

What good looks like

Grain or policy first, boxes second. Atomic publish. Safe reruns. Clear out-of-scope.

Scope control

Protect the critical path. Park optional marts after the SLA landing succeeds.

Interview framing details

You are expected to teach while you design. Start from a concrete failure, define the terms, then put the algorithm or layout on the board. Keep the scope tight: one grain, one SLA, one publication method.

Ask only the clarifying questions that would change your diagram. Write assumptions when the interviewer asks you to decide. Call out out-of-scope work so you do not burn the clock on tooling trivia.

A strong close restates the critical path, the failure mode you fear most, and what you would ship in week one.

Interview framing details

You are expected to teach while you design. Start from a concrete failure, define the terms, then put the algorithm or layout on the board. Keep the scope tight: one grain, one SLA, one publication method.

Ask only the clarifying questions that would change your diagram. Write assumptions when the interviewer asks you to decide. Call out out-of-scope work so you do not burn the clock on tooling trivia.

A strong close restates the critical path, the failure mode you fear most, and what you would ship in week one.