Copy-on-write (COW) and merge-on-read (MOR) are two different strategies for executing row-level updates and deletes in lakehouse table formats. Copy-on-write rewrites every Parquet data file containing an updated or deleted record, resulting in heavier write amplification and slower ingestion but guaranteeing optimal, fast read performance. Merge-on-read avoids rewriting full files by writing small supplemental delta files (such as delete files or deletion vectors) that readers reconcile at query time, prioritizing fast write throughput at the cost of slower read performance until compaction runs.
Trade-offs between write amplification and read speed
The fundamental trade-off balances write latency against read latency:
- In a Copy-on-write workflow, if an incoming batch modifies one single row inside a 500 MB Parquet file with two million rows, the engine must read the entire 500 MB file, modify that single row, and write a brand-new 500 MB Parquet file. For write-heavy pipelines with frequent CDC updates, this causes massive write amplification and high compute costs. However, downstream analytical queries read pure, uninterrupted Parquet files with zero merge overhead.
- In a Merge-on-read workflow, the engine does not rewrite the base file. Instead, it writes a small companion file. When a query scans the table, the query engine reads the base Parquet file and merges it on the fly with the delete or delta records, discarding deleted keys or projecting updated values.
How deletion files alter query execution
Different table formats implement these mechanisms under specific configurations:
- Apache Hudi explicitly exposes distinct table types: Copy on Write (storing pure Parquet files) and Merge on Read (storing base columnar files alongside row-based Avro log delta files).
- Apache Iceberg v2 uses merge-on-read with positional delete files or equality delete files, which readers reconcile into memory filters during scan phases.
- Delta Lake historically used COW for all updates, but introduced deletion vectors in recent versions to deliver MOR benefits without full file rewrites.
Choosing the right update strategy
Select COW when tables are read heavily by BI dashboards and updated infrequently in daily batches. Select MOR when streaming CDC pipelines ingest frequent micro-batch updates, scheduling background compaction jobs to merge delta files into clean Parquet files during off-peak windows.