Apache Beam is an open SDK and programming model for batch and streaming pipelines. You write a pipeline once (Python/Java) using Beam transforms (ParDo, GroupByKey, windows).
Cloud Dataflow is Google's fully managed runner for Beam. You submit a Beam pipeline; Dataflow handles workers, autoscaling, and stream/batch execution on GCP.
Beam pipeline code (Python/Java)
|
v
Dataflow runner (GCP)
|
+--> read Pub/Sub / GCS
+--> transform / window / join
+--> write BigQuery / GCSWhy the split matters
- Beam = portable API (can also run on Flink, Spark runners elsewhere)
- Dataflow = managed execution on GCP (ops-light)
Fresher mental model
Think of Beam as the *recipe* and Dataflow as the *kitchen that cooks it* on Google Cloud.
Typical DE jobs
- Streaming enrich of Pub/Sub events into BigQuery
- Large GCS → Parquet / BigQuery batch ETL
- Windowed aggregates (tumbling / session windows)
Interview tip: "Beam is the model; Dataflow is GCP's managed Beam runner." Contrast with writing raw Spark on Dataproc if asked.