Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. What is a data lake?

Cloud · Extra High-Value

What is a data lake?

Easycloud-24
data lakeobject storageparquetarchitecture

Question

What is a data lake?

Solution

A data lake is a central repository that stores large amounts of raw and processed data in open file formats on cheap object storage (S3 / GCS / ADLS). Schema can be applied when read; many engines share the same files.

Sources --> raw zone (JSON/CSV/Parquet)
              |
              v
         cleaned / curated zones
              |
              +--> Spark / Athena / BigQuery external / Databricks

Why lakes exist

Warehouses want structured tables; lakes keep *everything* cheaply for replay, ML, and multi-engine access. ELT patterns land raw first.

Healthy lake habits

  • Zones/layers (raw → curated)
  • Partitioned Parquet/Delta/Iceberg, not endless tiny files
  • Catalog + IAM governance (not a "data swamp")

Interview tip: "Lake = object storage + open formats + shared access." Contrast with dumping files with no ownership or schema as a swamp.

PreviousNext