Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. concurrent.futures ThreadPool vs ProcessPool

Python · Language Internals Interviewers Still Ask

concurrent.futures ThreadPool vs ProcessPool

Hardpython-68
concurrent-futuresthreadsprocessesgilio-bound

Question

How do ThreadPoolExecutor and ProcessPoolExecutor differ, and which would you use to download 10,000 files?

Solution

For downloading 10,000 files, use ThreadPoolExecutor. The work is waiting on the network, and threads can overlap those waits. ProcessPoolExecutor is for CPU-heavy work, such as parsing or compressing, where you need several cores running Python at once.

The GIL in one paragraph

CPython has a global interpreter lock, so only one thread runs Python bytecode at a time. For CPU-bound code, threads therefore give no speedup. But a thread waiting on I/O releases the lock, so other threads can run, and I/O-bound code does speed up with threads.

Threads for downloads

from concurrent.futures import ThreadPoolExecutor, as_completed
import requests

def download(url):
    r = requests.get(url, timeout=30)
    r.raise_for_status()
    return url, len(r.content)

with ThreadPoolExecutor(max_workers=32) as pool:
    futures = {pool.submit(download, u): u for u in urls}
    for fut in as_completed(futures):
        url = futures[fut]
        try:
            _, size = fut.result()
        except Exception as e:
            print("failed", url, e)       # log and keep going

Processes for CPU work

from concurrent.futures import ProcessPoolExecutor

with ProcessPoolExecutor(max_workers=8) as pool:
    parsed = list(pool.map(parse_file, paths, chunksize=10))

Each process has its own interpreter and memory, so they run truly in parallel. The cost: arguments and results are pickled and copied between processes, so passing huge objects is slow, and starting processes takes longer than starting threads. Functions must be defined at module level so they can be pickled.

Choosing max_workers

For I/O, more workers than cores is fine, often 16 to 64, limited by what the remote server and your network allow, since too many workers trigger rate limits. For CPU work, use about the number of cores (os.cpu_count()).

Handle failures

fut.result() re-raises any exception that happened in the worker, so wrap it in a try block, as above, or you lose the error until you call it. Decide whether one failure stops the run or is recorded and retried.

The mixed case

Download with threads, then parse with processes, in two stages. Say that for a very large fan-out, asyncio or a distributed framework such as Spark may fit better.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext