Problem
A Claude thread showed Provider turn start failed: ProviderAdapterSessionClosedError: claudeAgent adapter thread is closed (thread 67f4d809-8d85-48ea-95c0-068b5e368989, production release 5561b14f9c, 2026-10-02). The thread had looked stuck for about ten minutes. The error appeared only after Stop was pressed, so it reads as a crash when it is the result of the stop. The two messages sent during the stall were never queued.
What the logs show (all times UTC)
| Time |
Event |
| 15:19 |
Last provider event for the thread, then 18 minutes of silence (events per minute: 15:19 = 26, 15:20 to 15:37 = 0, 15:38 = 86) |
| 15:28:00 |
sendTurn #1 starts (model claude-sonnet-5-5, interaction mode default) and never returns |
| 15:38:20 |
sendTurn #2 starts and also never returns |
| 15:38:31.76 |
Both sendTurn spans fail with ProviderAdapterSessionClosedError; cause is Error: Query closed before response received |
| 15:38:31.76 to :33.77 |
interruptTurn runs stopSessionInternal, which takes 2006 ms |
| 15:38 onward |
New system/init events and the thread works again |
Sources: ~/.t3/userdata/logs/server.trace.ndjson.2 (the four sendTurn failure spans and the interruptTurn span) and ~/.t3/userdata/logs/provider/events.67f4d809-8d85-48ea-95c0-068b5e368989.log (event counts).
Root cause (probable)
Before sendTurn queues the message in apps/server/src/provider/Layers/ClaudeAdapter.ts, it awaits SDK control requests: query.setModel(...) when the model changed, and query.setPermissionMode(...) on every send that carries an interaction mode. These wait for the Claude CLI to answer and have no timeout. While the CLI is silent they block, and every later send on the thread queues behind the same stall.
interruptTurn then calls stopSessionInternal, whose query.close() rejects the pending request with Query closed before response received. toSessionError maps any message containing "closed" to ProviderAdapterSessionClosedError, which the orchestrator reports as a failed turn start. Stopping the thread on purpose is therefore reported as a provider failure.
Not verified
- Which control call was pending. The trace spans carry no attribute for it, so
setModel and setPermissionMode cannot be told apart.
- Why the CLI stopped emitting events at 15:19 (a long tool or subagent, or a hung process). The thread had background
Task subagents running (task_progress events, scaffold-dev).
- Whether upstream
main behaves the same. This is the same code path as upstream's, but only the fork's packaged build was observed.
Proposed fixes
- Bound the control requests in
sendTurn with a timeout, and skip setPermissionMode when the mode is unchanged (track it as currentApiModelId is tracked), so a busy CLI cannot hold a message.
- When the session was closed by our own
stopSessionInternal (context.stopped), report an interrupted or stopped outcome instead of a turn start failure. Narrow the includes("closed") match in toSessionError.
- Show a visible waiting state when
sendTurn has been pending for more than a few seconds, so the thread does not look dead.
- Add the pending control method to the
sendTurn span attributes so the next occurrence is diagnosable.
Severity
Medium. Stop recovers the thread, but messages sent during the stall are lost and the error text misleads. It was seen on the scaffold orchestrator thread, which runs long background subagents, so it may recur there.
Related
Upstream pingdotgg#2336 (closed) covers a different failure, a dangling resume cursor after the CLI is killed before its first turn. It is not this bug.
Reported from the production build; not yet reproduced on demand.
Problem
A Claude thread showed
Provider turn start failed: ProviderAdapterSessionClosedError: claudeAgent adapter thread is closed(thread67f4d809-8d85-48ea-95c0-068b5e368989, production release5561b14f9c, 2026-10-02). The thread had looked stuck for about ten minutes. The error appeared only after Stop was pressed, so it reads as a crash when it is the result of the stop. The two messages sent during the stall were never queued.What the logs show (all times UTC)
sendTurn#1 starts (modelclaude-sonnet-5-5, interaction modedefault) and never returnssendTurn#2 starts and also never returnssendTurnspans fail withProviderAdapterSessionClosedError; cause isError: Query closed before response receivedinterruptTurnrunsstopSessionInternal, which takes 2006 mssystem/initevents and the thread works againSources:
~/.t3/userdata/logs/server.trace.ndjson.2(the foursendTurnfailure spans and theinterruptTurnspan) and~/.t3/userdata/logs/provider/events.67f4d809-8d85-48ea-95c0-068b5e368989.log(event counts).Root cause (probable)
Before
sendTurnqueues the message inapps/server/src/provider/Layers/ClaudeAdapter.ts, it awaits SDK control requests:query.setModel(...)when the model changed, andquery.setPermissionMode(...)on every send that carries an interaction mode. These wait for the Claude CLI to answer and have no timeout. While the CLI is silent they block, and every later send on the thread queues behind the same stall.interruptTurnthen callsstopSessionInternal, whosequery.close()rejects the pending request withQuery closed before response received.toSessionErrormaps any message containing "closed" toProviderAdapterSessionClosedError, which the orchestrator reports as a failed turn start. Stopping the thread on purpose is therefore reported as a provider failure.Not verified
setModelandsetPermissionModecannot be told apart.Tasksubagents running (task_progressevents,scaffold-dev).mainbehaves the same. This is the same code path as upstream's, but only the fork's packaged build was observed.Proposed fixes
sendTurnwith a timeout, and skipsetPermissionModewhen the mode is unchanged (track it ascurrentApiModelIdis tracked), so a busy CLI cannot hold a message.stopSessionInternal(context.stopped), report an interrupted or stopped outcome instead of a turn start failure. Narrow theincludes("closed")match intoSessionError.sendTurnhas been pending for more than a few seconds, so the thread does not look dead.sendTurnspan attributes so the next occurrence is diagnosable.Severity
Medium. Stop recovers the thread, but messages sent during the stall are lost and the error text misleads. It was seen on the scaffold orchestrator thread, which runs long background subagents, so it may recur there.
Related
Upstream pingdotgg#2336 (closed) covers a different failure, a dangling resume cursor after the CLI is killed before its first turn. It is not this bug.
Reported from the production build; not yet reproduced on demand.