Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Write-audit-publish (WAP)

Data quality · Building Quality In

Write-audit-publish (WAP)

Harddata-quality-22
write-audit-publishapache-icebergdelta-lakeatomic-publishing

Question

What is the write-audit-publish pattern?

Solution

The write-audit-publish pattern writes pipeline output to an isolated staging area or branch, runs automated quality checks on that staged data, and publishes the records to production only if all checks pass. This pattern guarantees that downstream consumers never see corrupted rows, partial batch appends, or broken schemas during processing.

Mechanics of the staging branch

Traditional ETL pipelines append records directly into production tables and attempt rollbacks when tests fail, leaving consumers exposed to dirty reads. Write-audit-publish eliminates this operational hazard through a strict three-phase contract:

  • Write phase: The pipeline writes processed data into an isolated staging directory, shadow table, or metadata branch without touching the production catalog.
  • Audit phase: Automated assertions evaluate the isolated data against uniqueness rules, null constraints, referential relationships, and row volume expectations.
  • Publish phase: If all tests pass, the pipeline promotes the staged data into the production table using atomic metadata operations. If any test fails, the staged dataset is dropped or routed to an investigation workspace without altering production state.
Job Run -> [Write: Staging Branch] -> [Audit: Automated Quality Suite]
                                           |
                                  Pass     |     Fail
                                    v            v
                    [Publish: Atomic Promotion]  [Drop / Alert On-Call]

Modern lakehouses implement this workflow natively through multiple storage engines:

  • Apache Iceberg with Project Nessie or native branching commits writes to an isolated branch, runs assertions, and fast-forwards the main branch tag atomically.
  • Delta Lake writes into a staging location or shallow clone, validates data, and performs an atomic merge or partition overwrite into the target table.
  • BigQuery and Snowflake write to an intermediate table like staging_orders and execute an atomic metadata swap or partition copy command into production.

Consumer isolation and reliability

Because promotion happens as an atomic catalog metadata swap, analytical queries experience zero downtime and zero intermediate partial reads. Business intelligence dashboards and downstream ML feature stores query consistent snapshots with guaranteed quality boundaries.

PreviousNext