Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Cloud Dataflow and Apache Beam

Cloud · GCP

Cloud Dataflow and Apache Beam

Mediumcloud-08
gcpdataflowbeamstreamingbatch

Question

What are Cloud Dataflow and Apache Beam?

Solution

Apache Beam is an open SDK and programming model for batch and streaming pipelines. You write a pipeline once (Python/Java) using Beam transforms (ParDo, GroupByKey, windows).

Cloud Dataflow is Google's fully managed runner for Beam. You submit a Beam pipeline; Dataflow handles workers, autoscaling, and stream/batch execution on GCP.

Beam pipeline code (Python/Java)
        |
        v
  Dataflow runner (GCP)
        |
        +--> read Pub/Sub / GCS
        +--> transform / window / join
        +--> write BigQuery / GCS

Why the split matters

  • Beam = portable API (can also run on Flink, Spark runners elsewhere)
  • Dataflow = managed execution on GCP (ops-light)

Fresher mental model

Think of Beam as the *recipe* and Dataflow as the *kitchen that cooks it* on Google Cloud.

Typical DE jobs

  • Streaming enrich of Pub/Sub events into BigQuery
  • Large GCS → Parquet / BigQuery batch ETL
  • Windowed aggregates (tumbling / session windows)

Interview tip: "Beam is the model; Dataflow is GCP's managed Beam runner." Contrast with writing raw Spark on Dataproc if asked.

PreviousNext