Repository navigation
OAuthClientProvider auth lock is permanently poisoned when httpx closes async_auth_flow from a different task (RuntimeError: The current task is not holding this lock) #3382
Description
Activity
- addedv2Affects the v2 line (2.x on main)Affects the v2 line (2.x on main)v1Affects the v1.x maintenance lineAffects the v1.x maintenance line
on Aug 25, 2026 I reproduced the failure against current
main(4d6f87e8) and confirmed the proposed primitive change isolates the lifecycle problem:- A flow suspended inside
async with anyio.Lock()raisesRuntimeError: The current task is not holding this lockwhen a different asyncio task callsaclose(), and the lock remains held. - The same flow using
anyio.Semaphore(1)closes cleanly from the other task and the next flow acquires the semaphore. OAuthContext.lockis only used byOAuthClientProvider.async_auth_flow, which intentionally holds it across the generator's HTTP yields, so the change can stay local to the OAuth context and preserve the existingasync withserialization.
The regression should exercise the generator lifecycle rather than only checking the field type: advance one
async_auth_flowfrom one task until it yields, close it from a second task, then prove a subsequent flow can proceed. The test should use the repository's anyio style and bound the wait withfail_after, covering the cross-task teardown path without wall-clock assertions.This looks like a narrowly scoped compatibility fix rather than an API change. Please confirm whether
Semaphore(1)is the preferred direction; I will follow the repository's assignment/help-wanted gate before opening any PR.- A flow suspended inside
Production field report confirming this exact failure mode (mcp 2.0.0, CPython 3.11, macOS arm64, anyio 4.x asyncio backend).
A long-lived MCP client process (an ACP server hosting several OAuth streamable-HTTP servers: Notion, Composio, Consensus) hit this on 2026-08-22 03:44 and stayed poisoned for 4.5 days until the process was manually killed.
Trigger sequence (from logs):
GET stream disconnected, reconnecting in 1000ms...(server closed the standalone GET stream)- Session teardown closed the in-flight
async_auth_flowgenerator from a different task:
ERROR asyncio: Task exception was never retrieved future: <Task finished coro=<<async_generator_athrow without __name__>()> exception=RuntimeError('The current task is not holding this lock')> Traceback: File "mcp/client/auth/oauth2.py", line 601, in async_auth_flow response = yield request GeneratorExit During handling of the above exception, another exception occurred: File "mcp/client/auth/oauth2.py", line 582, in async_auth_flow async with self.context.lock: File "anyio/_core/_synchronization.py", line 166, in __aexit__ self.release() File "anyio/_backends/_asyncio.py", line 1854, in release raise RuntimeError("The current task is not holding this lock") RuntimeError: The current task is not holding this lockPost-poison behavior (matches your analysis exactly):
context.lockremained held forever. Every subsequent connect attempt for the affected servers hung insideasync with self.context.lock— producing zero network traffic (verified via socket inspection: no SYN, only CLOSE_WAIT zombies) — until each attempt hit its 300s connect timeout. Client-side symptom was a perfectly periodic retry metronome (300s backoff + 3×300s hung attempts = 25m07s cycle) that survived restarts of every other client process on the machine, because the poisoned state is in-memory per-process.Additional confirmation of the mechanism: the same OAuth servers handshaked fine (<1s initialize+list_tools) from fresh processes throughout the entire incident — tokens, network, and servers were never the problem. Only the poisoned process could not connect.
Killing the process fully resolved it. Strong +1 for the
Semaphore(1)direction (or any release path that doesn't assume the releasing task is the acquirer).I’d like to take this issue if it is still available.
I’ve reviewed the reproduction and the production report, and
the cross-task generator teardown behavior looks like the key
lifecycle issue.I can implement the Semaphore(1) change together with the
cross-task async_auth_flow regression test, following the
repository’s contribution workflow.If this is still available, please assign it to me.
This looks nasty to debug blind did the RuntimeError point you straight at the lock, or did you have to trace back from a hang/timeout first? Also curious if this showed up right after an SDK bump or you'd been running this version a while before it surfaced.
Kludex commented
on Oct 10, 2026 MemberMore actionsBoth reports identify
OAuthClientProvider.async_auth_flowholding ananyio.Lockacross a yield, which can fail on cross-task generator closure and leave the lock held. This is tracked in #2847, so I’m closing this as a duplicate. AI-assisted triage; I reviewed both reports.
Summary
OAuthClientProvider.async_auth_flowholdsself.context.lock(ananyio.Lock) across the entire httpx auth-flow generator, including everyyield(src/mcp/client/auth/oauth2.py:582in v2.0.0; same code is present on v2.1.0 and main).anyio.Lock.release()is bound to the acquiring task. When httpx closes the auth generator from a different task than the one that advanced it — which happens routinely when a request is cancelled mid-flight (network drop, timeout, task-group teardown) — theasync with__aexit__runs in the closing task and raises:The exception escapes into a fire-and-forget teardown task ("Task exception was never retrieved"), and the lock is left permanently held. Every subsequent request through the same
OAuthClientProviderthen blocks forever atasync with self.context.lock:without sending any HTTP. For a long-lived client that reuses the provider across reconnects, that server is dead until the whole process restarts.Environment
mcp2.0.0 (reproduced; the sameasync with self.context.lock:pattern is unchanged in 2.1.0 and on main)Observed traceback
Two captures from the same host, one per yield point (authorization-code exchange and the main
yield request):(The other capture is identical except the
GeneratorExitlands at line 746,token_response = yield await self._perform_authorization().)After this fires, every reconnect attempt for that server times out with no HTTP traffic — the flow generator never gets past line 582.
Minimal reproduction (no network needed)
Output on 2.0.0:
Root cause
anyio.Lockis task-bound by design; httpx makes no guarantee that the auth-flow generator is closed from the task that advanced it (cancellation/teardown commonly runsaclose()from a sibling task). Holding a task-bound lock across the generator's yields therefore poisons the lock on any cross-task close: the release both raises and never happens.Suggested fix
Serializing the auth flow per-context is still needed (token refresh must not race), but the primitive must allow release from the closing task. Options:
anyio.Semaphore(1)instead ofanyio.LockforOAuthContext.lock. anyio semaphores are not owner-bound, so release from the closing task is legal on both asyncio and trio backends. One-line change in theOAuthContextdataclass plus theasync withkeeps working.GeneratorExit/cross-task unwind releases via a tolerant path (catch the ownershipRuntimeErrorand force the lock back to a released state), though anyio has no public API for that today.Happy to send a PR for option 1 if that direction is acceptable.