What usually matters:
- Scheduler capacity: tune scheduler config / run HA schedulers
- Database: a properly sized production metadata DB
- DAG parsing: keep heavy business logic out of parse time
- Executor: distributed (Celery/K8s)
- Workers: scale horizontally
- Pools: protect limited shared resources
- DAG design: don't explode into unnecessary tasks
- Limits:
max_active_runs, task concurrency, scheduler parallelism
Interview tip: Scaling Airflow is mostly metadata DB + parse cost + executor/worker capacity + good DAG hygiene.