Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Data pipeline testing pyramid

Pipelines & scenarios · Core Concepts Not Yet Covered

Data pipeline testing pyramid

Mediumpipelines-59
testingdata-pipelinescidata-testsregression

Question

How do you test a data pipeline end to end?

Solution

Test a pipeline in layers, like a pyramid: many fast, small tests at the bottom, fewer slow and broad ones above, and checks that run on real data in production.

Layer 1: unit tests on transformation logic

Write the transformations as functions or SQL models with clear inputs and outputs. Test them with tiny hand-made data, including edge cases: NULLs, duplicates, empty input, the first and last day of a month, a negative amount. They run in seconds on a laptop or in CI, so developers run them all the time.

Layer 2: contract and schema tests at boundaries

Check that the data entering and leaving each stage has the expected columns, types and constraints. A source schema check at ingestion, and an output schema check on published tables, catch breaking changes between teams. Tools such as dbt tests, Great Expectations, or schema registries do this.

Layer 3: integration tests

Run the whole pipeline, or one major slice, on a small sample dataset in a dev environment, with real connections to storage and the orchestrator. They find problems that unit tests cannot: permissions, file formats, wrong paths, and ordering between steps. They are slower, so they run on pull requests or nightly.

Layer 4: data tests in production

On every real run, check the data itself: uniqueness and not-null on keys, accepted values, referential integrity, row counts within normal ranges, freshness, and totals reconciled to a source. These are your protection against bad inputs that no test could predict. Decide which failures block publishing and which only warn.

Layer 5: regression comparison

When changing logic, compare the new output with the old output on the same input, and review the differences. Data diff tools do this for tables: they show rows added, removed and changed, so a change that should affect 1 percent of rows does not quietly alter 40 percent.

CI gates

Pull requests run linting, unit tests, schema checks and a build of changed models against a dev schema or a clone. Merging is blocked if they fail. Deployment is automated, so what was tested is what ships.

What to say

No single test type is enough. Unit tests check your logic, contract tests check your neighbours, and production data tests check reality. Mention that you would start with tests on primary keys and row counts, because they catch most real incidents for little effort.

PreviousNext