Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Variables vs Connections vs XCom

Airflow & DAGs · Operating Airflow in Production

Variables vs Connections vs XCom

Easyairflow-58
variablesconnectionsxcomconfiguration

Question

What is the difference between Variables, Connections and XCom?

Solution

Variables, Connections, and XCom serve distinctly different purposes in Airflow. Variables hold shared configuration, Connections manage external system credentials, and XCom passes runtime metadata between tasks.

The three concepts

Each mechanism addresses a specific operational requirement:

  • Connections (airflow.models.Connection): Store connection parameters for external databases, cloud services, and APIs (such as hostname, port, login credentials, and extra JSON configs). They are referenced by Hooks and Operators, keep secrets masked in the UI, and can be retrieved from external secret managers.
  • Variables (airflow.models.Variable): Store global key-value configuration values (like default batch sizes, feature flags, or environment names). They are accessible across all DAGs and editable via the web UI.
  • XCom (Cross-Communication): Passes small runtime data between tasks within the same DAG run. When an upstream task returns a value, Airflow serializes it into the metadata database; a downstream task pulls that value using ti.xcom_pull().
# Pulling an XCom value produced by an upstream task
@task
def process_data(ti=None):
    file_path = ti.xcom_pull(task_ids="extract_task", key="output_path")
    print(f"Processing file: {file_path}")

The danger of large XCom payloads

By default, all three primitives store their data directly inside the relational metadata database (PostgreSQL or MySQL).

A severe trap is returning large datasets from Python tasks (such as pandas DataFrames, large JSON payloads, or multi-megabyte API responses). Serializing megabytes of data into the metadata database bloats database tables, fills WAL logs, exhausts database memory, and degrades scheduler query speed.

Never pass actual data records through XCom. Instead, write the raw data to cloud object storage (such as S3 or GCS) and push only the lightweight storage URI (e.g. s3://bucket/2025-01-01/orders.parquet) or row count to XCom. Alternatively, configure a custom XCom backend that automatically handles object store serialization under the hood.

PreviousNext