Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Full load vs incremental load

Batch & Streaming · Batch Design

Full load vs incremental load

Easybatch-streaming-48
batchincremental-loadfull-loaddata-pipelines

Question

When do you do a full reload instead of an incremental load?

Solution

A full table reload truncates and completely repopulates the destination dataset on every execution, making it the preferred approach for small tables, sources lacking dependable change timestamps, or tables experiencing untracked physical deletions. Incremental loading processes only rows modified or inserted since the previous watermark, which is mandatory for large datasets where reprocessing full historical tables exceeds compute budgets and batch windows. Full reloads are simple and self-healing, and teams frequently combine both patterns by running daily incremental loads alongside periodic full refreshes to correct data drift.

Scenarios where full reloads are preferred

Despite the popularity of incremental processing, full table reloads remain the best engineering choice in specific situations:

  • Small reference and lookup tables: Dimension tables like country codes, currency conversion tables, or internal product categories with fewer than a few hundred thousand rows can be truncated and reloaded in seconds.
  • Absence of reliable change tracking: If an upstream operational database lacks dependable updated_at timestamps, soft-delete flags, or change logs, incremental filtering cannot reliably identify modified records.
  • Untracked hard deletes: When upstream applications physically remove rows using SQL DELETE without recording tombstones, incremental queries based on timestamps will never discover that records were deleted.
  • Logic refactoring and disaster recovery: Full reloads are inherently idempotent and self-healing. If a transformation formula changes, rebuilding the entire table eliminates historical inconsistencies without complex backfills.

When incremental loading is mandatory

Incremental loads become necessary when dataset scale makes full reloads computationally impractical:

  • Tables containing millions or billions of rows, such as web telemetry, payment transaction ledgers, and sensor logs, cannot be queried and rewritten during a nightly batch window.
  • By tracking a high-watermark timestamp or consuming Change Data Capture (CDC) events, the pipeline extracts only the delta of records altered over the last hour or day, reducing compute costs and query runtimes.

The hybrid drift correction pattern

A battle-tested production pattern combines both techniques:

Execute fast hourly or daily incremental loads to keep dashboards fresh while minimizing warehouse costs. Concurrently, schedule a weekly or monthly full refresh during weekend maintenance windows. This periodic refresh automatically fixes subtle discrepancies, missed updates, or untracked hard deletes, ensuring the warehouse remains accurate over time.

PreviousNext