A data pipeline is the automated path that moves and shapes data from sources to destinations where people or systems can use it. It includes ingest, validate, transform, load, and observe.
End-to-end DE mental model
┌──────────┐ ingest ┌─────────────┐ transform ┌────────────┐ │ Sources │ -----------> │ Landing / │ -----------> │ Curated │ │ OLTP,API │ batch file │ Bronze raw │ Spark/dbt │ Silver/Gold│ │ events │ or CDC/stream│ + checks │ + tests │ marts │ └──────────┘ └─────────────┘ └─────┬──────┘ | v ┌────────────┐ │ Serve │ │ BI, ML, API│ └────────────┘ Cross-cutting: Airflow orchestration | quality tests | lineage | alerts | IAM
Batch pipeline example: nightly S3 files → Spark job → fct_orders → Tableau.
Streaming pipeline example: Kafka orders → Flink enrich → Redis + lake table within seconds.
Interview tip: Walk source → ingest → transform → serve, and name one reliability idea (idempotency, retries, data quality) without drowning in tools.