Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing

Cloud Platforms for Data Engineers

Progress0/14
x

What is the Cloud

  • What is cloud computing?12m
  • Cloud services overview12m
  • Data lakes versus warehouses12m
  • Cloud security basics12m

Cloud Mental ModelPreview

  • Why the cloud, and which one12m
  • IAM: who can do whatFree14m
  • Object storage as the bronze landing zone14m

Managed Data Services

  • Serverless SQL warehouses14m
  • Managed orchestration12m
  • Managed Spark12m
  • Serverless compute12m

Cost, Security, and Shipping It

  • Reading a cloud bill12m
  • Networking a data engineer actually needs12m
  • Capstone: deploy ingest to the cloud16m
Back to track
  1. Learn
  2. Cloud Platforms for Data Engineers
  3. What is the Cloud
  4. Data lakes versus warehouses

Lesson 3 of 14 · Theory first, then run it

Data lakes versus warehouses

cloudpythonbeginner12 min

Overview

A lake stores cheap files (bronze). A warehouse stores structured SQL tables. A lakehouse uses both in one design.

On this page7 sections›
  1. 1The decision
  2. 2What is at stake
  3. 3Option A vs Option B
  4. 4A worked comparison
  5. 5Common beginner questions
  6. 6What comes next
  7. 7Practice

The decision

A data lake is object storage full of files: JSON, CSV, Parquet, images, logs. You pay for bytes sitting in a bucket. You do not get a SQL table until a job or a warehouse reads those files. In the medallion pattern, bronze usually lands here: raw evidence, cheap to keep.

A data warehouse is structured tables you query with SQL. Columns have types. Analysts write SELECT. BigQuery, Redshift, and Synapse are the three names you will hear. You pay for the compute that scans those tables, not only for the bytes on disk.

A lakehouse keeps the cheap files and adds table-shaped access: Delta or Iceberg on the lake, or a warehouse that queries the files directly. Same bronze objects, SQL on top. This tab models the three words as a dict. It does not create a bucket or a dataset.

What is at stake

If you dump CSV into a warehouse as the only copy, storage gets expensive and you lose the original file when someone runs a bad UPDATE. If you only keep files in a lake, analysts cannot write a JOIN without a processing job first. Most companies do both: land bronze in the lake, publish gold as warehouse tables the dashboard trusts.

Job postings say 'data lake' and 'warehouse' in the same sentence. Mixing them up in an interview (calling S3 a warehouse, or calling BigQuery a lake) is a fast way to fail a screen. The words name different storage shapes.

Option A vs Option B

Think of a loading dock and a catalog. The lake is the dock: pallets (files) land cheap and messy. The warehouse is the catalog: labeled aisles (tables) you can query. A lakehouse is a dock that also has aisle labels on the pallets.

Files in the lake, tables in the warehouse
Bronze: raw files in the lakeSilver: cleaned files or tablesGold: warehouse tables for dashboards

Bronze is usually objects. Silver may stay as Parquet in the lake. Gold is often warehouse tables.

Data Lake architecture showing raw ingestion into decoupled object storage and compute layers
Data Lake decoupled architecture: raw events land in scalable object storage before distributed compute engines execute ETL and analytical queries.
Source: Wikimedia CommonsCC BY-SA 4.0

Object storage (S3, Cloud Storage, Blob Storage) addresses files by bucket plus key. A prefix such as bronze/orders/dt=2026-08-26/ is an aisle label, not a nested folder the disk created. A warehouse table has a schema: order_id STRING, gmv NUMERIC. You do not SELECT * from a random JSON blob until something parses it.

Pick the shape from the consumer. Dashboards want tables. Replay and audit want the original files.

ShapeWhat you storeHow you queryTypical cost instinct
LakeFiles (JSON, CSV, Parquet)Spark, pandas, or a warehouse external tableBytes stored; cheap bronze
WarehouseTyped tablesSQL (SELECT, JOIN, GROUP BY)Bytes scanned or warehouse compute
LakehouseFiles plus table metadataSQL on the lake, or both toolsLake storage plus some compute

Medallion is a quality ladder, not a product. Bronze preserves the source. Silver is cleaned, typed, deduplicated. Gold is the grain the business named (daily GMV, one row per order). You can build that ladder in a lake, in a warehouse, or across both. Production bronze is still usually files.

Trace the storage shapes

Watch the same BlueCart orders land as a file, become a trusted silver row, then a gold metric. Predict lake versus warehouse before the next panel.

A worked comparison

Name the three storage shapes in a dict. The Python editor cannot list a real bucket. It can make the vocabulary checkable.

PythonThree words, three values
WHERE = {
    "lake": "files",
    "warehouse": "tables",
    "lakehouse": "both",
}
print("bronze lands as", WHERE["lake"])
print("analysts query", WHERE["warehouse"])
print("a lakehouse is", WHERE["lakehouse"])

A tiny inventory shows why both exist. The lake holds the raw file. The warehouse holds the table an analyst can JOIN. Dropping either copy loses a job: replay, or a dashboard.

PythonBronze file, gold table
INVENTORY = [
    {"layer": "bronze", "place": "lake", "object": "orders/dt=2026-08-26/part.json"},
    {"layer": "gold", "place": "warehouse", "object": "gold.daily_gmv"},
]
for row in INVENTORY:
    print(row["layer"], "in the", row["place"], "as", row["object"])

Common beginner questions

Is a lake just a cheaper warehouse?

No. A warehouse enforces a schema on write or on query and is built for SQL. A lake accepts any file. Cheap is true for cold bytes. Cheap is false if every analyst must write Spark to answer 'yesterday's GMV.'

Do I need a lakehouse product?

Not on day one. Landing files in Cloud Storage or S3 and loading BigQuery or Redshift is already a lake plus a warehouse. Lakehouse formats (Delta, Iceberg) matter when many jobs need ACID tables on the same files.

Where does pandas fit?

pandas is compute on one machine, often reading a lake file or a warehouse extract. It is not the lake and not the warehouse. The Pandas track is the wrangling. This lesson is where those frames live in production.

A folder of CSVs is not a warehouse

If the only 'schema' is the header row, you have a lake (or a share drive). A warehouse is tables with types, permissions, and a query engine.

PySpark already named medallion

If you completed PySpark, bronze / silver / gold is the same ladder. Here the cloud question is: which layer is files, and which layer is SQL tables.

What comes next

The next lesson covers cloud security basics: IAM and least privilege, encryption at rest, and keeping secrets out of git. A lake full of customer files with a public bucket is an incident, not a design.

Practice

Run Sample to print the three storage shapes. Then complete Exercise: assign {"lake": "files", "warehouse": "tables", "lakehouse": "both"} to result and print it.

Practicals · load into the editor

After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.

Rate:
Was this useful?
Cloud services overviewCloud security basics