Speculative execution starts a second copy of a task that is running much slower than its siblings. Whichever copy finishes first wins, and the other is killed. It is meant for slow machines, not slow data.
How it works
Spark watches the tasks in a stage. Once most of them have finished, it compares each running task against the median duration of the finished ones. A task far above that gets a duplicate on another executor. It is off by default and you turn it on with spark.speculation=true. The related settings (multiplier, quantile, interval) control how slow is "slow" and how often Spark checks.
When it helps
A bad disk, a noisy neighbour VM, or a flaky network link makes one task slow even though its data is normal. Running the same task elsewhere finishes quickly. This happens now and then on big clusters and cloud spot nodes.
When it does not help
Data skew. If one task has 50 times more rows than the others, the copy on a different executor has the same 50 times more rows and is slow too. Both copies run for an hour and you have wasted a core. The fix for skew is to change the data layout or the join, as in the skew question.
When it can hurt
Tasks with side effects. Say a foreachPartition sends rows to an external database or an HTTP API. With speculation, two copies of the same task may both send, and you get duplicates. Writes to file outputs are protected by the output committer, but writes to external systems are not. Speculation is only safe if tasks are idempotent. It also uses extra resources, which matters on a shared cluster.
A good short answer: it fixes slow nodes, not slow data, and should be paired with idempotent writes.