Skip to content

ProcessPoolExecutor.shutdown(wait=True) deadlocks when worker exit messages fill the result pipe #158570

Description

@tfzee

Bug report

ProcessPoolExecutor.shutdown(wait=True) can hang forever when the pool has many workers.

When shutting down, _ExecutorManagerThread._join_executor_internals() sends one None sentinel per worker and then calls p.join() on every worker. Each worker handles its sentinel by calling result_queue.put(os.getpid()) in _process_worker before it exits. By then the manager thread has stopped reading result_queue, so nothing drains it while it waits in join().
Once the result pipe is full, the next exiting worker blocks in write() while holding the queue's write lock. All remaining workers then block on that lock, and join() never returns.
Each exit message is about 21 bytes, so a default 64 KiB pipe only fills with roughly 3000 workers.
On Linux, however an unprivileged user who is over fs.pipe-user-pages-soft gets new pipes at the minimum size only 1/2 Pages. This is low enough to be high on high core count machines that run many processes.

I found this when running LLVM's lit test runner on a 512-thread machine. All tests finished, then shutdown() hung with 464 workers. At that point 389 workers had exited, one was blocked in pipe_write, and the rest were waiting on the queue's write lock.

Reproducer. F_SETPIPE_SZ is only used to simulate the small-pipe fallback deterministically:

import concurrent.futures
import fcntl
import multiprocessing

F_SETPIPE_SZ = 1031

def init(barrier):
    global _barrier
    _barrier = barrier

def work(_):
    _barrier.wait()  # every worker must be running a task at the same time

if __name__ == "__main__":
    n = 400
    ex = concurrent.futures.ProcessPoolExecutor(
        max_workers=n, initializer=init,
        initargs=(multiprocessing.Barrier(n),))
    fcntl.fcntl(ex._result_queue._writer.fileno(), F_SETPIPE_SZ, 4096)
    list(ex.map(work, range(n)))  # all n workers are now alive
    ex.shutdown(wait=True)        # hangs
    print("done")

With n = 100 this finishes in about 0.03s. With n = 400 it hangs on 3.10.11, 3.12.3, 3.13.15, 3.14.7 and main. Without the F_SETPIPE_SZ call, it also hangs when run by a user who is over pipe-user-pages-soft.

Stack of the manager thread while hung (main):

File ".../multiprocessing/process.py", line 156 in join
File ".../concurrent/futures/process.py", line 684 in _join_executor_internals
File ".../concurrent/futures/process.py", line 667 in join_executor_internals

Proposed fix: keep draining the result queue while joining the workers in _join_executor_internals(). When that point is reached there is also no work items left as such we can drain them safely.
I have a patch that fixes it with a regression test and will open a PR shortly.

The analysis and reproducer were prepared with help from an AI assistant (Claude). I reviewed and verified them.

CPython versions tested on:

3.10, 3.12, 3.13, 3.14, CPython main branch

Operating systems tested on:

Linux

Linked PRs

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    stdlibStandard Library Python modules in the Lib/ directorytopic-multiprocessingtype-bugAn unexpected behavior, bug, or error

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions