spark-submit is the command that sends your application to a cluster manager and starts it. It takes your script or JAR plus options that say where to run it and how many resources to ask for.
A typical call
spark-submit \ --master yarn \ --deploy-mode cluster \ --num-executors 20 \ --executor-cores 5 \ --executor-memory 18g \ --driver-memory 4g \ --conf spark.sql.shuffle.partitions=400 \ --py-files deps.zip \ etl_orders.py --run-date 2025-03-01
Options you will set most often
--master: which cluster manager (yarn,k8s://...,spark://host:7077, orlocal[*]for testing).--deploy-mode: client or cluster.--num-executors,--executor-cores,--executor-memory: the size and number of workers. With dynamic allocation on, you set min and max through--confinstead.--driver-memory: raise it if you collect big results or broadcast large tables.--conf key=value: any Spark setting.--py-filesand--archives: ship Python modules, or a packed virtualenv, so executors can import your code. A common mistake is installing a library only on the driver and then seeingModuleNotFoundErrorinside a UDF on the executors.
Anything after the script name goes to your own program as arguments.
Which setting wins
The same setting can come from three places. From highest to lowest priority: values set in code on SparkConf or SparkSession.builder, then flags on spark-submit, then spark-defaults.conf. So a hard-coded .config("spark.executor.memory", ...) in the script overrides the flag you pass on the command line, which is a classic source of "why is my setting ignored".
Some settings, such as driver memory in client mode, must be given before the JVM starts, so setting them in code is too late. Pass them at submit time.