Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Data lakehouse in practice

Data platform · Platform Architecture

Data lakehouse in practice

Mediumdata-platform-35
lakehouseapache-icebergdelta-laketable-formats

Question

What actually makes something a lakehouse rather than a data lake?

Solution

A data lakehouse is an architecture that combines the cost-effective scalability of object storage with the transactional guarantees, schema governance, and query performance of a traditional data warehouse. What transforms a raw data lake into a true lakehouse is an open table format (such as Apache Iceberg or Delta Lake) coupled with a centralized catalog that manages metadata, enforces ACID transactions, and enables schema evolution. Without a table format managing file state, you simply have a directory of static files susceptible to partial writes and inconsistent reads.

Metadata catalogs and ACID guarantees

In a traditional data lake, compute engines write raw Parquet or ORC files directly to cloud object storage folders like S3 or GCS. Because object storage lacks file locking and atomic multi-file updates, failed pipeline jobs leave orphaned files, and readers experience dirty reads during write operations.

A lakehouse resolves this through layered architectural tiers:

  • Open file formats: Underlying records are stored in compressed columnar formats like Apache Parquet.
  • Open table formats: Formats like Apache Iceberg, Delta Lake, and Apache Hudi track explicit file manifests and snapshot logs. Writes commit via atomic snapshot swaps, providing full ACID isolation, time-travel queries, and automated partition evolution.
  • Central catalog and governance: Catalogs such as Unity Catalog, Polaris, or Nessie maintain single-source access controls, column masking, and data lineage across engines.
Raw Data Lake:   S3 Bucket -> Folders/Files (No transactions, dirty reads)
Lakehouse Stack: S3 Storage -> Parquet Files -> Table Format (Iceberg) -> Catalog (Polaris)
                                                       ^
                                (ACID Snapshots, Schema Enforcement, Time Travel)

A single source of truth powers both analytics and machine learning:

Storage layers and analytical engines

The primary operational advantage of a lakehouse is compute engine decoupling. Multiple disparate engines query the exact same physical copy of data without requiring ETL replication into proprietary database storage:

  • SQL analytical engines: Trino, StarRocks, and Snowflake query Iceberg tables directly for business intelligence dashboards.
  • Machine learning frameworks: PySpark, Ray, and TensorFlow read the same underlying Parquet files for feature extraction and model training without moving data across network boundaries.

Implementing table formats like Iceberg or Delta Lake prevents vendor lock-in while providing the reliability and performance historically exclusive to monolithic data warehouses.

PreviousNext