Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Delta change data feed

File formats & storage · Table Formats in Depth

Delta change data feed

Mediumfile-formats-37
delta-lakecdcchange-data-feedstreamingmedallion

Question

What is Delta Lake's Change Data Feed?

Solution

Delta Lake's Change Data Feed (CDF) is a table capability that records row-level change events, capturing inserts, deletes, and updates across table versions. When enabled, Delta logs change types alongside pre-image and post-image values for every modified record, allowing downstream pipelines to consume only the incremental row mutations between two table versions instead of rescanning entire tables. This feature powers efficient medallion architecture pipelines, propagating incremental updates from silver to gold aggregation tables without expensive full-table reconciliations.

Tracking row-level change events

Table Update (Version 5 -> 6):
  Change Data Feed Output:
  _change_type        id   status     amount
  ------------------------------------------
  update_preimage     42   pending    100.00   (Row state before commit)
  update_postimage    42   completed  100.00   (Row state after commit)
  insert              99   active      50.00   (Newly added record)
  delete              17   cancelled   20.00   (Deleted record)

CDF captures four distinct mutation event types in its changelog stream:

  • insert: Represents brand-new records added to the table.
  • delete: Represents rows removed from the table.
  • update_preimage: Captures the exact state of a row before an UPDATE or MERGE modified it, allowing downstream streaming consumers to subtract previous measures from rolling metrics.
  • update_postimage: Captures the updated state of the row immediately after the transaction committed.

Propagating changes across medallion layers

In a typical lakehouse medallion architecture, raw events arrive into bronze tables and merge into cleansed silver tables:

  • Updating downstream gold aggregate tables (such as daily revenue per merchant) from silver usually requires running full daily recalculations if changes include historical corrections.
  • With CDF enabled on the silver table, a Structured Streaming job reads the change feed, consuming only the net deltas to adjust running aggregates incrementally in gold.
  • This pattern eliminates the need to scan multi-terabyte silver tables on every pipeline run.

Incremental consumption patterns

In the Apache Iceberg ecosystem, equivalent functionality is achieved through incremental scan reads and changelog views (such as Iceberg's table.changes() table function). Enabling CDF introduces a minor write overhead because Delta must write small change files alongside main data files during updates and deletes, but downstream compute savings in incremental processing vastly outweigh the modest write cost.

PreviousNext