Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Speculative execution

PySpark · Execution Model

Speculative execution

Mediumpyspark-48
speculative-executionstragglersidempotency

Question

What is speculative execution, and when can it hurt?

Solution

Speculative execution starts a second copy of a task that is running much slower than its siblings. Whichever copy finishes first wins, and the other is killed. It is meant for slow machines, not slow data.

How it works

Spark watches the tasks in a stage. Once most of them have finished, it compares each running task against the median duration of the finished ones. A task far above that gets a duplicate on another executor. It is off by default and you turn it on with spark.speculation=true. The related settings (multiplier, quantile, interval) control how slow is "slow" and how often Spark checks.

When it helps

A bad disk, a noisy neighbour VM, or a flaky network link makes one task slow even though its data is normal. Running the same task elsewhere finishes quickly. This happens now and then on big clusters and cloud spot nodes.

When it does not help

Data skew. If one task has 50 times more rows than the others, the copy on a different executor has the same 50 times more rows and is slow too. Both copies run for an hour and you have wasted a core. The fix for skew is to change the data layout or the join, as in the skew question.

When it can hurt

Tasks with side effects. Say a foreachPartition sends rows to an external database or an HTTP API. With speculation, two copies of the same task may both send, and you get duplicates. Writes to file outputs are protected by the output committer, but writes to external systems are not. Speculation is only safe if tasks are idempotent. It also uses extra resources, which matters on a shared cluster.

A good short answer: it fixes slow nodes, not slow data, and should be paired with idempotent writes.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext