Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is incremental processing?

Pipelines & scenarios · Core Pipeline Concepts

What is incremental processing?

Mediumpipe-07
incrementalwatermarkfull refreshmergecost

Question

What is incremental processing, and how does it differ from a full refresh?

Solution

Incremental processing only reads and writes new or changed data since the last successful run. A full refresh rebuilds the whole target from scratch every time.

Full refresh each night:
  scan ALL history -> rewrite entire table

Incremental:
  watermark = last_success_ts
  read source WHERE updated_at > watermark
  merge into target
  advance watermark

Common incremental strategies

1. Timestamp watermark on updated_at / ingestion_time 2. High-water mark on an auto-increment id 3. CDC change stream 4. Partition append for immutable daily files (dt=YYYY-MM-DD)

Why incrementals win

Faster, cheaper, less load on sources. Needed when history is terabytes.

Why full refresh still exists

Small dimensions, logic rewrites, or when incremental bugs (late updates, clock skew) make full rebuild safer.

Late update trap

A row updated yesterday with updated_at backdated can be missed. CDC or periodic full reconciliations help.

Interview tip: "Incremental = process the delta; full refresh = rebuild everything." Mention watermarks and merge, plus the late-update risk.

PreviousNext