Overview
A lake stores cheap files (bronze). A warehouse stores structured SQL tables. A lakehouse uses both in one design.
On this page7 sections
The decision
A data lake is object storage full of files: JSON, CSV, Parquet, images, logs. You pay for bytes sitting in a bucket. You do not get a SQL table until a job or a warehouse reads those files. In the medallion pattern, bronze usually lands here: raw evidence, cheap to keep.
A data warehouse is structured tables you query with SQL. Columns have types. Analysts write SELECT. BigQuery, Redshift, and Synapse are the three names you will hear. You pay for the compute that scans those tables, not only for the bytes on disk.
A lakehouse keeps the cheap files and adds table-shaped access: Delta or Iceberg on the lake, or a warehouse that queries the files directly. Same bronze objects, SQL on top. This tab models the three words as a dict. It does not create a bucket or a dataset.
What is at stake
If you dump CSV into a warehouse as the only copy, storage gets expensive and you lose the original file when someone runs a bad UPDATE. If you only keep files in a lake, analysts cannot write a JOIN without a processing job first. Most companies do both: land bronze in the lake, publish gold as warehouse tables the dashboard trusts.
Job postings say 'data lake' and 'warehouse' in the same sentence. Mixing them up in an interview (calling S3 a warehouse, or calling BigQuery a lake) is a fast way to fail a screen. The words name different storage shapes.
Option A vs Option B
Think of a loading dock and a catalog. The lake is the dock: pallets (files) land cheap and messy. The warehouse is the catalog: labeled aisles (tables) you can query. A lakehouse is a dock that also has aisle labels on the pallets.
Bronze is usually objects. Silver may stay as Parquet in the lake. Gold is often warehouse tables.

Object storage (S3, Cloud Storage, Blob Storage) addresses files by bucket plus key. A prefix such as bronze/orders/dt=2026-08-26/ is an aisle label, not a nested folder the disk created. A warehouse table has a schema: order_id STRING, gmv NUMERIC. You do not SELECT * from a random JSON blob until something parses it.
Pick the shape from the consumer. Dashboards want tables. Replay and audit want the original files.
| Shape | What you store | How you query | Typical cost instinct |
|---|---|---|---|
| Lake | Files (JSON, CSV, Parquet) | Spark, pandas, or a warehouse external table | Bytes stored; cheap bronze |
| Warehouse | Typed tables | SQL (SELECT, JOIN, GROUP BY) | Bytes scanned or warehouse compute |
| Lakehouse | Files plus table metadata | SQL on the lake, or both tools | Lake storage plus some compute |
Medallion is a quality ladder, not a product. Bronze preserves the source. Silver is cleaned, typed, deduplicated. Gold is the grain the business named (daily GMV, one row per order). You can build that ladder in a lake, in a warehouse, or across both. Production bronze is still usually files.
Trace the storage shapes
Watch the same BlueCart orders land as a file, become a trusted silver row, then a gold metric. Predict lake versus warehouse before the next panel.
A worked comparison
Name the three storage shapes in a dict. The Python editor cannot list a real bucket. It can make the vocabulary checkable.
WHERE = {
"lake": "files",
"warehouse": "tables",
"lakehouse": "both",
}
print("bronze lands as", WHERE["lake"])
print("analysts query", WHERE["warehouse"])
print("a lakehouse is", WHERE["lakehouse"])A tiny inventory shows why both exist. The lake holds the raw file. The warehouse holds the table an analyst can JOIN. Dropping either copy loses a job: replay, or a dashboard.
INVENTORY = [
{"layer": "bronze", "place": "lake", "object": "orders/dt=2026-08-26/part.json"},
{"layer": "gold", "place": "warehouse", "object": "gold.daily_gmv"},
]
for row in INVENTORY:
print(row["layer"], "in the", row["place"], "as", row["object"])Common beginner questions
Is a lake just a cheaper warehouse?
No. A warehouse enforces a schema on write or on query and is built for SQL. A lake accepts any file. Cheap is true for cold bytes. Cheap is false if every analyst must write Spark to answer 'yesterday's GMV.'
Do I need a lakehouse product?
Not on day one. Landing files in Cloud Storage or S3 and loading BigQuery or Redshift is already a lake plus a warehouse. Lakehouse formats (Delta, Iceberg) matter when many jobs need ACID tables on the same files.
Where does pandas fit?
pandas is compute on one machine, often reading a lake file or a warehouse extract. It is not the lake and not the warehouse. The Pandas track is the wrangling. This lesson is where those frames live in production.
A folder of CSVs is not a warehouse
If the only 'schema' is the header row, you have a lake (or a share drive). A warehouse is tables with types, permissions, and a query engine.
PySpark already named medallion
If you completed PySpark, bronze / silver / gold is the same ladder. Here the cloud question is: which layer is files, and which layer is SQL tables.
What comes next
The next lesson covers cloud security basics: IAM and least privilege, encryption at rest, and keeping secrets out of git. A lake full of customer files with a public bucket is an incident, not a design.
Practice
Run Sample to print the three storage shapes. Then complete Exercise: assign {"lake": "files", "warehouse": "tables", "lakehouse": "both"} to result and print it.
Practicals · load into the editor
After you read the theory, run these in the pane on the right. They execute in this tab, no cluster.