Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Backfill without hurting production

Airflow & DAGs · Operating Airflow in Production

Backfill without hurting production

Hardairflow-64
scenariobackfillcapacity-planningproduction-safetyidempotency

Question

You need to backfill 90 days of a heavy DAG. How do you do it without hurting production or the source system?

Solution

Executing a 90-day backfill for a heavy pipeline requires disciplined pacing, strict idempotency verification, and clear resource throttling to avoid degrading production workloads and database health.

Step 1: Verify idempotency and parameterization

Before queuing dozens of runs, verify that every task in the pipeline is strictly idempotent:

  • Check that SQL queries filter on data_interval_start and data_interval_end, not real-time functions like NOW().
  • Confirm that data write operations use partition overwrites or upserts (MERGE), rather than blind INSERT INTO statements that generate duplicate rows if retried.
  • Configure tasks to write to date-partitioned storage locations (such as s3://bucket/year=2024/month=10/day=05/) so historical writes do not touch current live partitions.

Step 2: Pilot run on a single day

Never launch a 90-day backfill all at once. Trigger a pilot backfill for 1 to 2 historical days:

# Testing a single historical interval in Airflow CLI
airflow dags backfill heavy_pipeline --start-date 2024-10-01 --end-date 2024-10-02

Inspect the pilot run in the web UI. Verify that row counts match expectations, runtime matches historical benchmarks, and downstream data warehouse tables report correct metrics.

Step 3: Concurrency throttling and pool allocation

If you launch 90 daily runs simultaneously without guards, Airflow will attempt to execute hundreds of tasks in parallel. This will saturate database connection pools, exhaust worker memory, and cause warehouse query queuing that delays production pipelines:

  • Restrict concurrent DAG runs by temporarily setting max_active_runs = 2 or max_active_runs = 3 in the DAG definition.
  • Assign database extraction tasks to a dedicated Airflow pool (e.g. backfill_pool, limited to 2 slots) so backfill queries never consume production database slots.
  • Schedule the backfill execution during off-peak hours (such as overnight or over a weekend) when live reporting dashboards and transactional databases experience lower traffic.

Step 4: Execution, monitoring, and communication

In Airflow 3, backfills are first-class operations managed directly in the React web UI or API, allowing you to pause, inspect, and cancel batches easily.

Notify downstream analytics and business intelligence stakeholders before starting, so they know historical partitions are being recomputed. Monitor cloud warehouse compute credits closely to ensure the backfill does not exhaust monthly spend quotas.

PreviousNext