A data lake stores large amounts of raw and semi-structured data cheaply as files (Parquet/ORC/JSON on S3/GCS/ADLS). A data warehouse stores structured, query-optimized tables for analytics (Snowflake, BigQuery, Redshift).
lake: files + schemas optional, flexible, cheap scale warehouse: tables + SQL governance, fast BI, stronger structure
Typical use
- Lake: land everything, data science, large historical archives, Spark jobs
- Warehouse: governed marts, BI dashboards, documented metrics
Lakehouse idea
Table formats (Iceberg/Delta/Hudi) give warehouse-like tables on lake storage: ACID, time travel, schema evolution.
Interview tip: Do not claim one replaces the other. Many companies use a lake for raw/history and a warehouse or lakehouse SQL engine for consumption.