Both are ways to share data between the driver and executors without putting it into the DataFrame. A broadcast variable sends a read-only value from the driver to every executor once. An accumulator goes the other way: tasks add to it, and the driver reads the total.
Broadcast variables
Use one when many tasks need the same lookup data, such as a dictionary of country codes:
codes = {"IN": "India", "US": "United States"}
b = spark.sparkContext.broadcast(codes)
rdd.map(lambda r: (r[0], b.value.get(r[1])))Without a broadcast, Python would ship the dictionary inside every task's closure. With a broadcast, each executor gets one copy and tasks share it. It only helps while the value is small enough for each executor's memory.
Do not confuse this with a broadcast join. A broadcast join (F.broadcast(df)) is an optimizer strategy where the small DataFrame is copied to all executors to avoid a shuffle. It uses the same mechanism underneath, but you do not manage it as a variable.
Accumulators
An accumulator is a counter, often used for bookkeeping such as "how many rows had a bad date":
bad = spark.sparkContext.accumulator(0)
def check(row):
if row.order_date is None:
bad.add(1)
df.foreach(check)
print(bad.value)Tasks can only add. Only the driver can read the value.
The reliability catch
If an accumulator is updated inside a transformation such as map, and a task is retried or a stage is recomputed, the update can be applied again, so you count too much. Spark guarantees exactly-once updates only for accumulators updated inside actions such as foreach. And an accumulator in a transformation that never runs (because the result was not used) never updates at all. So treat accumulators in transformations as rough metrics, not as numbers to base decisions on.