Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. One Big Table (OBT)

Data modeling · Modern Modeling Approaches

One Big Table (OBT)

Mediumdata-modeling-49
one-big-tablestar-schemacolumnar-storagedata-marts

Question

What is the One Big Table approach, and when is it better than a star schema?

Solution

The One Big Table (OBT) approach pre-joins fact and dimension tables into a single, wide, fully denormalized table for analytics. It excels in modern columnar cloud warehouses like BigQuery and Snowflake because columnar compression and partition pruning minimize storage overhead and eliminate join latency for end-user dashboards.

How One Big Table works

In an OBT architecture, an analyst querying customer purchase history queries a single table with fifty or one hundred columns instead of writing multiple joins across fact_orders, dim_customer, dim_store, and dim_date. Because columnar warehouses only scan the specific columns referenced in a query, wide tables do not penalize queries that only touch three fields. Eliminating runtime joins also avoids data shuffles across worker nodes.

Star schema comparison

Despite query speed benefits, OBT has real architectural trade-offs:

  • Data duplication: Customer addresses and product categories repeat across millions of rows, increasing storage footprint and write times.
  • Slow SCD management: Updating a customer attribute requires rewriting millions of transaction rows rather than updating a single record in a dimension table.
  • Schema evolution friction: Adding or renaming an attribute ripples across a massive table, requiring extensive backfills.

The hybrid architecture

The industry standard pattern avoids picking one exclusively. Model the core warehouse as a Kimball star schema in the silver or transformation layer to maintain auditability, conformed dimensions, and clean surrogate keys. Then, build OBT marts on top in the gold layer using automated dbt models. This gives BI users simple, performant single-table querying without sacrificing data integrity.

PreviousNext