Apache Hudi (Hadoop Upserts Deletes and Incrementals) is a lakehouse table format aimed at incremental data and upsert-heavy pipelines: CDC landing, near-real-time tables, and efficient upserts on cloud storage.
Hudi table types (know these names)
- Copy-on-Write (CoW): updates rewrite Parquet files; reads stay simple/fast
- Merge-on-Read (MoR): base files + log/delta files; merge at read (or compact later) for faster writes
CDC / Kafka
-> Hudi upserts (CoW rewrite OR MoR log append)
-> queries see latest view after merge/compactionComparison table (Delta vs Iceberg vs Hudi)
| Capability | Delta Lake | Iceberg | Hudi | |-------------------------|-------------------------|----------------------------|-------------------------------| | Core idea | Tx log on Parquet | Snapshots + manifests | Upsert/incremental first | | ACID / time travel | Yes | Yes | Yes | | Upsert / CDC focus | Strong MERGE | Strong, engine-dependent | Core design goal | | Write patterns | Batch + streaming merges| Batch + streaming writers | CoW / MoR trade-offs | | Compaction | OPTIMIZE / vacuum | rewrite / expire snapshots | Compaction critical for MoR | | Multi-engine | Growing | Very strong | Strong Spark/Flink story |
Interview tip: Frame Hudi as "built for incremental upserts and MoR/CoW trade-offs," Delta as "transaction log lakehouse," Iceberg as "open metadata + multi-engine." Pick based on CDC intensity and engine mix, not brand loyalty.