A data lakehouse combines lake-style cheap object storage with warehouse-like table features: ACID transactions, schema enforcement, time travel, and SQL performance, usually via formats like Delta Lake, Apache Iceberg, or Apache Hudi.
Classic split:
lake (files, flexible, cheap) | warehouse (tables, governed SQL)
Lakehouse:
object storage (S3/GCS/ADLS)
+
table format (Iceberg/Delta/Hudi)
+
engines (Spark, Trino, warehouse engines)
=
one copy of data, many engines, ACID tablesWhat you gain
- Keep raw and curated data on the lake
- Concurrent readers/writers with commits (not "folder of Parquet chaos")
- Time travel / rollback after bad writes
- Schema evolution with clearer contracts
Interview tip: "Lakehouse = lake storage + transactional table format + SQL engines." Name Iceberg/Delta/Hudi and medallion layers if asked for architecture.