Repository navigation
asyncio: SelectorEventLoop busy-loops at 100% CPU forever when the self-pipe socketpair reaches EOF #156344
Description
Activity
Reproduced and confirmed on a real Windows box (3.13.15 and a self-built main): after the pair hits EOF, a 3-second idle sleep() loop fires _read_from_self 582,692 times (1.23s CPU) — one core fully pinned, loop otherwise functional but permanently spinning.
Proposed fix in #156345, mirroring the proactor-side approach from #156343: _read_from_self rebuilds the socketpair on clean EOF instead of leaving the reader registered on the dead socket. Selector-specific extra care: the process-wide signal.set_wakeup_fd registration is migrated only when it actually points at the loop's own _csock (a foreign registration is restored untouched), and from a worker thread the old write end is kept open so signal delivery survives (set_wakeup_fd is main-thread-only).
After the patch: 1 callback (the rebuild itself), 0.00s CPU over the same 3s idle, cross-thread call_soon_threadsafe still wakes the loop. Full test_asyncio suite green on Windows (34/34 files) and macOS; six new regression tests, all failing on the pre-fix code.
Same trigger reproduces on the selector loop too, and rebuild-on-EOF fixes it there as well — details in #156333 to avoid duplicating.
Tracking this over in #156333 — the WFP filter-driver trigger hits the selector loop the same way (clean EOF, no event logged anywhere), and it kills every idle asyncio process on the box at the same instant, both loop types included. That thread is now the field-data thread for both issues; the rebuild-on-EOF fix here (#156345) already survived it on a real machine.
The field-data archive is up in #156333 (pangi's comment with cpython-156333-field-data.zip, and the discussion under it). Directory 02 is the selector-relevant exhibit: the spin survived an explicit WindowsSelectorEventLoopPolicy switch on the same machine, samples rotating through _select / _process_events / _read_from_self — so the EOF-busy-loop is not proactor-specific behavior, and the rebuild-on-EOF fix here addresses the same root cause. Post-fix control (08) covers both loop types, all rebuilds EOF, zero runaways in three days.
Metadata
Metadata
Assignees
Labels
Projects
- StatusShow more project fieldsTodo
Bug description
The selector-side twin of #156333.
BaseSelectorEventLoopalso wakes itself through a self-socket created bysocket.socketpair(). On Windows this is a loopback TCP pair, and if the connection reaches a clean EOF while the loop runs,_read_from_selfreturns but the reader is still registered on the dead fd —select()reports it readable immediately, forever, pinning one core at 100% with nothing logged.Lib/asyncio/selector_events.py:At EOF,
recv()returnsb'', the loopbreaks, and the callback returns. But the fd is still registered via_add_reader(self._ssock.fileno(), self._read_from_self)from_make_self_pipe(). A closed-for-read socket is permanently "readable", so everyselect()iteration re-fires the callback — a tight spin. The loop is otherwise idle and never recovers.The proactor half of this bug is fixed by #156343 (rebuild the pair on the empty result). The same OS teardown of the loopback pair that triggers it triggers this on any
WindowsSelectorEventLoopPolicyprocess.Reproducer (deterministic, Windows)
Measured on Windows 11, CPython main (3.16.0a0, self-built): 582,692 callback invocations, 1.23s CPU, during a 3-second idle sleep. Expected ~0.
Note the instrumentation must patch the class before the loop is created —
_make_self_pipe()captures the bound method at registration time, so patching an instance afterwards does not intercept the already-armed reader (which is also why the spin is invisible to profilers that hook late).Environment
_read_from_selfcode path shown above is unchanged on mainNotes on scope
shutdown(SHUT_WR)works on any platform, though.ConnectionResetErrorfrom the pending recv instead of a clean EOF; that path is currently unhandled too (noted in gh-156333: rebuild the proactor self-pipe on EOF instead of busy-looping #156343).Suggested direction
On
recv() == b'', rebuild the self-pipe the way #156343 does for the proactor loop: allocate the new pair first, move any signal wakeup fd registration (set_wakeup_fd()returns the previous fd, so it can be probed and moved only when it names the old socket — Unix loops with signal handlers), remove the old reader, close the old sockets, and register the reader on the new fd.Linked PRs