A data lake is a central repository that stores large amounts of raw and processed data in open file formats on cheap object storage (S3 / GCS / ADLS). Schema can be applied when read; many engines share the same files.
Sources --> raw zone (JSON/CSV/Parquet)
|
v
cleaned / curated zones
|
+--> Spark / Athena / BigQuery external / DatabricksWhy lakes exist
Warehouses want structured tables; lakes keep *everything* cheaply for replay, ML, and multi-engine access. ELT patterns land raw first.
Healthy lake habits
- Zones/layers (raw → curated)
- Partitioned Parquet/Delta/Iceberg, not endless tiny files
- Catalog + IAM governance (not a "data swamp")
Interview tip: "Lake = object storage + open formats + shared access." Contrast with dumping files with no ownership or schema as a swamp.