Spark jobs run as a coordinated distributed system with three main roles.
Components
+------------------+
| Cluster Manager |
| (YARN/K8s/Stand) |
+--------+---------+
| allocates
v
+-------------+ +------+------+ +-------------+
| Driver |---->| Executor 1 | | Executor 2 |
| (your app) | | tasks+cache | ... | tasks+cache |
+-------------+ +-------------+ +-------------+1. Driver: runs your main program, builds the logical/physical plan, schedules tasks, tracks progress, and collects small results. 2. Cluster Manager: grants CPU/memory (YARN, Kubernetes, or Spark Standalone). 3. Executors: JVM processes on worker nodes that run tasks, store shuffle data, and hold cached partitions.
What to say in an interview
- The driver is the brain; executors are the workers.
- If the driver dies, the application usually fails.
- Executor failure is recoverable for recomputable RDDs/DataFrames (lineage).
- Keep the driver light: do not
collect()huge datasets to the driver.
Practical implication
# Bad: pulls all rows to the driver
rows = df.collect()
# Better: aggregate on executors, bring back a small result
summary = df.groupBy("country").count().collect()