Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Environments: dev, staging, prod for data

Pipelines & scenarios · Core Concepts Not Yet Covered

Environments: dev, staging, prod for data

Mediumpipelines-60
environmentsdevstagingprodci-cd

Question

How do you set up dev, staging and prod environments for data pipelines?

Solution

Keep development, staging and production separated, so mistakes in one cannot damage the next, and promote code through them automatically. The data is the hard part, because developers need realistic data without getting access to everything in production.

Separate the environments

Use separate projects, accounts or at least separate databases and schemas for each environment (dev_asha, staging, prod). Each has its own compute, its own storage paths, and its own credentials and service accounts. A dev job should not have permission to write to prod tables at all, so a wrong config cannot overwrite them.

Data for development

  • Dev works on a sample or a subset of production data, or a recent partition range, which is quick and cheap.
  • Zero-copy clones (Snowflake, BigQuery, Delta) can give a full-size copy of production tables in seconds, at a small storage cost, for testing. Clone with care for sensitive data.
  • Mask or tokenize personal data in non-production environments, or use synthetic data, since developer laptops and test systems are less protected.
  • Each developer can have their own schema, so work does not collide.

Staging mirrors production

Staging uses the same configuration, same code path and similar data volumes (or a fraction with the same shape) as production, so that problems appear before release. It is where integration tests and performance checks run, with a copy of production schedules.

Promotion by CI/CD

feature branch -> PR checks (tests, lint) -> merge -> deploy to staging -> tests -> approval -> deploy to prod

The same artifact (the same code version) moves through each step, with only the environment configuration changing, such as the catalog name, bucket and secrets. Infrastructure as code (Terraform, Databricks Asset Bundles) makes environments consistent.

Secrets and access

Each environment has its own secrets in a secret manager, and nothing is hard-coded in the repo. Production data access is limited to service accounts and a small group, with audit logs. People debug production with read-only access, and changes go through the pipeline.

Avoid testing in prod

If the only place to see realistic behaviour is prod, people will change things in prod. Make staging good enough that this is not needed. If a hotfix in prod is unavoidable, follow it with the same fix through the normal path.

PreviousNext