A data pipeline is the end-to-end path that moves data from sources to consumers, with transforms, quality checks, schedules, and ownership along the way.
It is not just "a Spark job." It is the whole flow: extract, land, clean, model, serve, monitor.
Big mental model (keep this diagram in your head)
┌─────────────┐ ┌──────────────┐ ┌─────────────────┐ │ Sources │ │ Ingestion │ │ Storage layers │ │ OLTP DBs │──▶│ batch/CDC │──▶│ lake / WH / │ │ APIs/SaaS │ │ files/Kafka │ │ lakehouse │ │ events │ └──────────────┘ └────────┬────────┘ └─────────────┘ │ v ┌──────────────────────────────────┐ │ Transform & model │ │ Spark / SQL / dbt / pandas │ │ bronze → silver → gold (or │ │ raw → staging → marts) │ └────────────────┬─────────────────┘ │ ┌──────────────────────┼──────────────────────┐ v v v ┌──────────┐ ┌──────────┐ ┌──────────┐ │ BI / SQL │ │ ML / │ │ Apps / │ │ dashboards│ │ features │ │ APIs │ └──────────┘ └──────────┘ └──────────┘ Cross-cutting (every serious pipeline has these): Orchestration (Airflow/Dagster) · Idempotent writes · DQ tests Monitoring / SLAs · Security (IAM, secrets) · Lineage / ownership
Pipeline vs DAG
- Pipeline = business data flow (what data moves and why).
- DAG = orchestration graph (how tasks are scheduled and depended).
Fresher example
Nightly ecommerce pipeline: pull orders from Postgres → land Parquet in S3 → Spark cleans → dbt builds fct_orders → Looker dashboard refreshes.
Interview tip: Start with sources → land → transform → serve, then name orchestration, quality, and monitoring. That diagram alone shows you think in systems, not scripts.