Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is a data lakehouse?

Pipelines & scenarios · Design & Scenarios

What is a data lakehouse?

Mediumpipe-17
lakehouseDeltaIcebergHudiACID

Question

What is a data lakehouse?

Solution

A data lakehouse combines lake-style cheap object storage with warehouse-like table features: ACID transactions, schema enforcement, time travel, and SQL performance, usually via formats like Delta Lake, Apache Iceberg, or Apache Hudi.

Classic split:
  lake (files, flexible, cheap)  |  warehouse (tables, governed SQL)

Lakehouse:
  object storage (S3/GCS/ADLS)
        +
  table format (Iceberg/Delta/Hudi)
        +
  engines (Spark, Trino, warehouse engines)
        =
  one copy of data, many engines, ACID tables

What you gain

  • Keep raw and curated data on the lake
  • Concurrent readers/writers with commits (not "folder of Parquet chaos")
  • Time travel / rollback after bad writes
  • Schema evolution with clearer contracts

Interview tip: "Lakehouse = lake storage + transactional table format + SQL engines." Name Iceberg/Delta/Hudi and medallion layers if asked for architecture.

PreviousNext