An executor is a JVM process launched on a worker node for the lifetime of a Spark application. It runs tasks, stores shuffle data, and can hold cached partitions.
Responsibilities
Executor |- task threads (cores) |- memory regions (execution + storage) |- disk for spill / shuffle / cache overflow |- reports metrics/status to the driver
How it relates to resources
- cores: parallel task slots in that executor
- memory: heap (and sometimes off-heap) for execution and cache
- Multiple executors per node are common depending on cluster config
# Common submit-time settings (cluster dependent) # --executor-memory 8g # --executor-cores 4 # --num-executors 50
Failure behavior
If an executor dies, Spark reschedules its tasks on other executors using lineage (for recomputable data). Cached blocks on the dead executor are lost and may be recomputed.
Interview tip
Contrast driver (coordinates) vs executor (computes). Mention that Python UDFs also involve Python workers alongside the JVM executor.