Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. The modern data stack

Data platform · Platform Architecture

The modern data stack

Easydata-platform-34
modern-data-stackplatform-architecturedbtdata-warehousing

Question

What does "modern data stack" mean, and what are its weaknesses?

Solution

The term modern data stack refers to a cloud-native, modular data architecture built around a central cloud data warehouse, managed SaaS ingestion connectors, dbt for SQL transformations, git-backed orchestrators, and browser-based BI tools. It gained popularity because it allowed small teams to stand up an operational analytics platform within days using standard SQL. However, its modularity created severe pain points including vendor sprawl, fragmented governance, compounding subscription costs, and poor support for real-time streaming and machine learning workloads.

Composition of modular stacks

The core paradigm of the modern data stack decoupled the data lifecycle into best-of-breed specialized SaaS layers:

  • Managed ingestion: Tools like Fivetran and Airbyte handle EL (extract and load) connectors, streaming source data into raw warehouse schemas without custom code.
  • Cloud storage and compute: Snowflake, Google BigQuery, or Databricks serve as the centralized storage and execution engine.
  • Transformation layer: dbt (data build tool) enables analytics engineers to write modular SQL select statements with built-in version control and automated testing.
  • Orchestration and observability: Tools like Dagster or Prefect coordinate schedules, while Monte Carlo or Elementary monitor data anomalies and freshness.
  • Business intelligence: Looker, Metabase, or Preset provide semantic layers and self-serve dashboard exploration.
SaaS Sources -> Ingestion (Fivetran) -> Warehouse (Snowflake) -> Transform (dbt) -> BI (Looker)
                     ^                         ^                     ^
             [Point Vendor 1]          [Point Vendor 2]      [Point Vendor 3]

These decoupled layers introduce significant operational friction:

Operational debt and tool consolidation

While modular tools accelerate initial deployment, maintaining five to eight separate SaaS vendors creates architectural headaches:

  • Tool sprawl and billing overhead: Each vendor charges separate licensing fees, minimum seat costs, or consumption markups, causing platform expenditures to spiral unpredictably.
  • Fragmented governance: Access policies defined in Snowflake do not sync cleanly to BI dashboards or ingestion connectors, requiring redundant role definitions across multiple interfaces.
  • Streaming and ML limitations: The modern data stack is fundamentally batch-oriented and SQL-centric, making it ill-suited for microsecond event processing or feature engineering pipelines.
  • Industry consolidation: Modern engineering teams increasingly favor unified platforms like Databricks or Snowflake that incorporate ingestion, transformation, governance, and ML capabilities under a single administrative roof.
PreviousNext