Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Design an ingestion framework for 200 sources

Pipelines & scenarios · System Design Questions

Design an ingestion framework for 200 sources

Hardpipelines-34
scenarioingestion-frameworkmetadata-drivenconfigconnectors

Question

Your company has 200 data sources (APIs, databases, files). How do you design ingestion so you don't write 200 pipelines?

Solution

Do not write 200 pipelines. Write a small number of generic ones, driven by configuration, so that adding a source means adding a config entry. The code is shared, and the differences between sources live in data.

The config

Each source has a record, kept in Git or a control table:

- name: crm_accounts
  type: postgres
  connection: crm_replica
  table: public.accounts
  load: incremental
  watermark_column: updated_at
  primary_key: [account_id]
  schedule: "0 * * * *"
  owner: sales-data@company.com
  freshness_sla_minutes: 90

A scheduler (such as Airflow generating one DAG per config entry, or a single DAG with dynamic task mapping) reads the config and runs the matching generic ingestion routine.

Reuse connectors

For common sources, buy or adopt a connector tool such as Fivetran or Airbyte instead of writing code. Write custom connectors only for odd sources. All connectors, bought or built, should land data the same way: the same raw format (say Parquet), the same path pattern (raw/<source>/<table>/load_date=...), and the same technical columns (_ingested_at, _source, _batch_id).

Common services around it

  • A generic quality step: row count, null primary key, duplicate key, freshness. The thresholds come from config.
  • Schema handling: detect drift, allow additive columns, alert on breaking changes.
  • Retries, backoff and rate limiting for APIs, once, in shared code.
  • A run log table recording what ran, how many rows, how long, and status.

Ownership and visibility

Each source has an owner who gets alerts, and a dashboard shows every source with its last success and freshness. If 12 sources are late, you see it on one screen.

Onboarding

A new source should be a pull request that adds a config entry (and credentials in the secret manager). If it needs code, you built the framework too narrowly. Review the exceptions regularly and fold the common ones into the framework.

Trade-offs

A framework is itself software that needs owners, tests and documentation. For 200 sources it pays off. For 10, a tool like Airbyte and some scripts is enough. State that you would start with the three or four source types that cover most of the 200, and extend after.

PreviousNext