Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. System Design

System design interview

Lakehouse vs Traditional Warehouse

MediumPro55 min read

Compare Iceberg/Delta/Hudi lakehouses vs Snowflake/BigQuery/Redshift on latency, ACID, cost, openness, and ML.

system-designlakehousewarehouseicebergdelta

Interview framing

Compare lakehouse formats (Iceberg, Delta Lake, Hudi) with traditional warehouses (Snowflake, BigQuery, Redshift). Discuss query latency, ACID, cost models, openness, and ML compatibility.

Background from first principles

A traditional cloud warehouse stores analytical tables in a managed service. You load data, SQL works well, governance and BI UX are polished, and you pay a productized compute/storage model. Many companies succeed here.

A lakehouse puts warehouse-like table semantics on cheap object storage using open table formats. Files are typically Parquet. Metadata layers (Iceberg/Delta/Hudi) provide transactions, schema evolution, and time travel so multiple engines (Spark, Trino, sometimes warehouses themselves) can share one copy of the facts.

The pain that creates lakehouses: copying the same sales facts into five systems for BI, ML, and ad hoc Spark, then watching definitions drift.

Without sharing:
  raw --> warehouse copy
      --> ML copy
      --> yet another export

With lakehouse:
  object storage tables (open format)
      ^
      +-- Spark / Trino / WH readers

Full problem statement

Give a practical comparison for a mid-size DE org that already has cloud object storage and some warehouse spend. Cover open table formats, managed warehouse storage/compute, ACID at table level, ML and BI access, cost levers, operational burden (small files, compaction), and when hybrid is normal. Avoid vendor wars; be accurate.

What "good" looks like

You name the duplication pain. You define lakehouse clearly. You compare axes: latency, ACID, cost, openness, ML. You admit warehouses often win out-of-the-box BI latency. You admit lakehouses need performance engineering. You recommend a hybrid pattern many companies actually run.

Clarifying questions

  • Primary consumers: BI, ML, or both?
  • Existing Snowflake/BQ commitment and skills?
  • Need for multi-engine access to the same tables?
  • Who will own compaction and table maintenance?
  • Regulatory / governance tooling expectations?

Scale prompts

  • Storage cost of second and third copies.
  • Compaction frequency for streaming small files.
  • BI query latency targets vs ML scan throughput.
  • Engineering time cost vs managed service premium.

Out of scope

  • Claiming lakehouses need no warehouse-like discipline.
  • Absolutism that one side makes the other obsolete.
  • Ignoring governance on the lake.

First principles expanded

Analytical data needs:

1. Durable storage. 2. A table abstraction (schema, snapshots, transactions). 3. Compute engines to query and transform. 4. Governance for people and tools.

Traditional warehouses bundle many of these. Lakehouses assemble open storage + open table format + chosen engines. Neither absolves you from modeling, quality, and access control.

Restated problem

"Compare Iceberg/Delta/Hudi lakehouses with Snowflake/BigQuery/Redshift-style warehouses on latency, ACID, cost, openness, and ML. Recommend a practical pattern for a mid-size DE org that already has object storage and some warehouse spend."

What good looks like

  • Duplication pain story.
  • Correct definition of open table formats.
  • Balanced comparison (no fanboying).
  • Small-file ops honesty.
  • Hybrid recommendation.

Clarifying questions expanded

  • BI vs ML traffic split?
  • Contractual warehouse commit?
  • Need Spark and SQL on same tables?
  • Platform staffing for compaction?
  • Governance tool preferences?

Estimation prompts

  • Cost of extra full copies per year.
  • Engineering hours for table maintenance.
  • BI latency targets (p95 seconds).
  • Streaming file creation rate and compaction interval.

Out of scope

  • Declaring warehouses obsolete.
  • Declaring lakehouses free of ops.
  • Ignoring security on object storage.

Room expectations and out-of-scope polish

Interviewers want you to sound like someone who has paid a cloud bill and compacted tables, not someone reciting a vendor slide. Mention governance: object storage buckets still need IAM, row/column controls, and PII handling. Mention that "open" does not mean "unguarded."

Keep multi-region DR and exotic catalog federation out of scope unless asked. Keep the conversation on copies, engines, latency, cost, and ops ownership.