Interview framing
Compare lakehouse formats (Iceberg, Delta Lake, Hudi) with traditional warehouses (Snowflake, BigQuery, Redshift). Discuss query latency, ACID, cost models, openness, and ML compatibility.
Background from first principles
A traditional cloud warehouse stores analytical tables in a managed service. You load data, SQL works well, governance and BI UX are polished, and you pay a productized compute/storage model. Many companies succeed here.
A lakehouse puts warehouse-like table semantics on cheap object storage using open table formats. Files are typically Parquet. Metadata layers (Iceberg/Delta/Hudi) provide transactions, schema evolution, and time travel so multiple engines (Spark, Trino, sometimes warehouses themselves) can share one copy of the facts.
The pain that creates lakehouses: copying the same sales facts into five systems for BI, ML, and ad hoc Spark, then watching definitions drift.
Without sharing:
raw --> warehouse copy
--> ML copy
--> yet another export
With lakehouse:
object storage tables (open format)
^
+-- Spark / Trino / WH readersFull problem statement
Give a practical comparison for a mid-size DE org that already has cloud object storage and some warehouse spend. Cover open table formats, managed warehouse storage/compute, ACID at table level, ML and BI access, cost levers, operational burden (small files, compaction), and when hybrid is normal. Avoid vendor wars; be accurate.
What "good" looks like
You name the duplication pain. You define lakehouse clearly. You compare axes: latency, ACID, cost, openness, ML. You admit warehouses often win out-of-the-box BI latency. You admit lakehouses need performance engineering. You recommend a hybrid pattern many companies actually run.
Clarifying questions
- Primary consumers: BI, ML, or both?
- Existing Snowflake/BQ commitment and skills?
- Need for multi-engine access to the same tables?
- Who will own compaction and table maintenance?
- Regulatory / governance tooling expectations?
Scale prompts
- Storage cost of second and third copies.
- Compaction frequency for streaming small files.
- BI query latency targets vs ML scan throughput.
- Engineering time cost vs managed service premium.
Out of scope
- Claiming lakehouses need no warehouse-like discipline.
- Absolutism that one side makes the other obsolete.
- Ignoring governance on the lake.
First principles expanded
Analytical data needs:
1. Durable storage. 2. A table abstraction (schema, snapshots, transactions). 3. Compute engines to query and transform. 4. Governance for people and tools.
Traditional warehouses bundle many of these. Lakehouses assemble open storage + open table format + chosen engines. Neither absolves you from modeling, quality, and access control.
Restated problem
"Compare Iceberg/Delta/Hudi lakehouses with Snowflake/BigQuery/Redshift-style warehouses on latency, ACID, cost, openness, and ML. Recommend a practical pattern for a mid-size DE org that already has object storage and some warehouse spend."
What good looks like
- Duplication pain story.
- Correct definition of open table formats.
- Balanced comparison (no fanboying).
- Small-file ops honesty.
- Hybrid recommendation.
Clarifying questions expanded
- BI vs ML traffic split?
- Contractual warehouse commit?
- Need Spark and SQL on same tables?
- Platform staffing for compaction?
- Governance tool preferences?
Estimation prompts
- Cost of extra full copies per year.
- Engineering hours for table maintenance.
- BI latency targets (p95 seconds).
- Streaming file creation rate and compaction interval.
Out of scope
- Declaring warehouses obsolete.
- Declaring lakehouses free of ops.
- Ignoring security on object storage.
Room expectations and out-of-scope polish
Interviewers want you to sound like someone who has paid a cloud bill and compacted tables, not someone reciting a vendor slide. Mention governance: object storage buckets still need IAM, row/column controls, and PII handling. Mention that "open" does not mean "unguarded."
Keep multi-region DR and exotic catalog federation out of scope unless asked. Keep the conversation on copies, engines, latency, cost, and ops ownership.