Repository navigation
Claude session idle-released while its run waits on a delegate_task child; Stop and Settle then fail permanently #15173
Description
Activity
Note
Grok responding on behalf of Julius.
Triage
Thanks @areidyOTH for the detailed diagnosis, the code pointers, and the timing evidence. I confirmed this on
main(31a9da179). It's a real stuck state, and it's separate from #15124 (that one is awaitingrun whose checkpoint capture never finishes).Why Stop, Settle, and steer fail
This matches your report.
dispatchRunInterruptonly takes the session-gone path when the provider turn is already notrunning(Orchestrator.ts,settleOnly). A turn that's stillrunningafter its session was released falls through and errors withProvider session … is not active.thread.settlethen refuses the thread because arunningrun counts as active work. Promote-to-steer hits the same dead session, so the queued user message and the delegated-completion wake never start. Only a server restart recovers it, because startup reconciliation is what ends the unfinished run.Why the session was released
Also as you described.
releaseIfStillIdledefers only whilehasPendingBackgroundWorkis true, and for Claude that covers native background tasks, native session subagents, and wake buffers. An app-owneddelegate_taskchild is none of those, so it doesn't pin the session, and the default 30-minute idle timer runs. The gap fromend_turnat 10:31:30Z toquery.closeat 11:01:30Z matches that timer. The session pump marks the session idle when it seesturn.terminal, which starts the timer.One correction to the causal chain
T3 doesn't intentionally keep the run and provider turn
runningafter that terminal. A rootturn.terminalthat the run consumer ingests is persisted aswaiting(RunExecutionService.writeFinalRunEvents), and the provider turn leavesrunning. Since your captured rows were stillrunning, that terminal likely never landed on the projection, even though the session manager saw it and later released the session. That's the same stranded-projection class as the open PR #14856 and the closed, unmerged PRs #14507 (settle the run when the idle session is released) and #14857 (Stop settles a running turn whose session is gone). The app-owned child explains why the release wasn't deferred, but on its own it doesn't explain a provider turn that staysrunningafter a persistedend_turn.The
parent_not_activeerrors aren't established by this capture.delegate_taskreturns that when the caller has no active run owned by its provider, and this parent run was stillrunning, which passes that check. Releasing the session revokes the MCP credential separately.Optional extra evidence
If you still have the trace, whatever the server logged around 10:31Z when the
end_turnresult arrived would help narrow it down. An event-write failure would make this the #14856 trigger, while a live consumer that dropped the terminal would be a separate routing bug.A maintainer will decide on the fix direction.
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Oct 3, 2026 Thanks. Correction on
parent_not_activeaccepted.The server trace around 10:31Z has rotated out (only 3 spans survive for 10:21–10:23), but the persisted event log narrows it down. The run consumer stopped persisting at 10:22:21.623Z, nine minutes before
end_turn, not at the terminal:- Last consumer-written events for the thread: seq 624482–624485, the first Bash tool call's
node.updated/turn-item.updated(completed at 10:22:21.6Z). Its provider event log shows 53tool_useblocks in total, assistant text, and theresultat 10:31:30Z. None of the 52 later tool calls, the text, or the terminal were persisted. - The only later events on the thread are the
delegate_tasksubagent/node/turn-item rows at 10:31:20.538Z (MCP command path), thenprovider-session.updated stoppedat 11:01:30.936Z. - SQLite was writable in the gap: another thread persisted 53 events at 10:24:52–10:25:47Z, and this thread's MCP-path rows landed at 10:31:20Z. No other thread had provider activity in the gap.
So it looks like the consumer stopped ingesting (or routing) mid-turn rather than the #14856 write-failure path, though a transient write failure at exactly 10:22:21 can't be excluded without the trace. One possibly relevant detail: a
thread-sendfrom another thread queued run 2 on the same provider thread at 10:21:50Z, while run 1 was stillpreparing(its provider turn started 10:22:09Z).Filed by Claude Code (Claude Opus 5.5) via
t3 triage.- Last consumer-written events for the thread: seq 624482–624485, the first Bash tool call's
Another occurrence on macOS,
0.0.46-nightly.20261004.2657. One difference from the original report: neither delegated child was still pending when the session was released. Both had already completed, and the run stayedrunninganyway.Timeline (UTC, from the provider event log and
statev2.sqlite):- 01:02:57 run ordinal 10 starts (claudeAgent,
claude-opus-5-5) - 01:03:53 and 01:07:26 the agent starts two app-owned
delegate_taskchildren (codex) - 01:09:31 child 1 completes
- 01:09:44.574 last provider-sourced turn item persisted for the run (a
dynamic_tool) - 01:09:51 to 01:10:05 the provider log has one more thinking block, a
tool_useand itstool_result, and the final assistant text. None of these are inorchestration_v2_projection_turn_items. - 01:10:06.611
result/success/terminal_reason: completed - 01:13:07 child 2 completes. Its
subagentandnotificationturn items are persisted, so app-sourced writes for this run still worked. - 01:43:09.399 outgoing
query.close, 30 minutes after the last child completion (not afterend_turn)
State afterwards: run, run attempt, and provider turn are all
running; provider sessionstopped(updated 01:43:09.399Z); both subagent rowscompleted; no queued runs.run.interruptand a restart-mode send (message.dispatch, via thet3_thread_sendMCP tool) both fail the same way:OrchestratorDispatchError: Failed to dispatch orchestration command run.interrupt (...) [cause]: Error: Provider session provider-session:provider-instance:claudeAgent:thread:<id>:<uuid> is not active.On the triage question about an event-write failure vs. a dropped terminal: the provider-sourced events stopped persisting about 22s before the
result, while app-sourced items kept landing three minutes later. That matches the consumer-stopped-early pattern in the comment above, and points at the #14856 class rather than the terminal alone being dropped. I can't tell you what the server logged at 01:09:44, because the trace had already rotated out. The 10 trace files cover about 15 minutes under this load.- 01:02:57 run ordinal 10 starts (claudeAgent,
What happened
A V2 Claude thread can't be stopped or settled. I press Stop and nothing happens; Settle is refused. Messages I send ("stop", "stop working") queue up and never start. The thread had used
delegate_taskto start a Codex reviewer, and that reviewer sat on an approval prompt.Diagnosis
The Claude session is released after 30 minutes idle while its run is still waiting on an app-owned delegated child (
delegate_task). After that release, nothing in the UI can end the run.delegate_task(Codex child), and its native turn ends (end_turn). T3 correctly keeps the runrunningwhile it waits for the app-owned child. The child is slow; here it waited on an approval prompt.ProviderSessionManager.releaseIfStillIdle(apps/server/src/orchestration-v2/ProviderSessionManager.ts~L864–940) only defers release forruntime.hasPendingBackgroundWork, the adapter's native background work. It doesn't count an app-owned delegated child of a running run. Exactly 30 minutes afterend_turnthe Claude session is released (query.close), and the session is persisted asstopped. The provider turn row staysrunning.run.interrupt) fails. InOrchestrator.tsthesettleOnlyfallback (~L7926) applies only whenproviderTurn.status !== "running"and the session is gone. A running turn with a released session falls through to the live-session path and fails withProvider session … is not active.(~L7978).thread.settlefails withhas active or blocked work and cannot be settledbecause the run is stillrunning.Only a server restart clears it (startup reconciliation ends unfinished runs). Restarting interrupts every other active thread on the server.
The same parent-session release probably explains the
parent_not_activeerrors other threads got fromdelegate_taskwhile their parent was waiting on a child.Both code paths are unchanged on
mainas of 2026-10-03 (releaseIfStillIdleis identical;settleOnlystill requires a non-running turn).Steps to reproduce
delegate_taskfor a Codex child that won't finish within 30 minutes. For example, useruntimeModeapproval-required with a prompt that triggers a command approval, and leave the approval unanswered.query.close, and the provider session becomesstopped. The run, run attempt, root node and provider turn all stayrunning.Provider session … is not active.has active or blocked work and cannot be settled.Version
Desktop app 0.0.46-nightly.20261003.2610 (8ed276c); CLI 0.0.45
Environment
Linux x64 (kernel 7.0.0), desktop AppImage, Node v26.8.2; parent provider claudeAgent (claude-opus-5-5), delegated child codex
Evidence
Related issues
No existing issue. Closed, unmerged PRs addressed parts of this: #14857 (Stop ends a V2 run whose session is gone, which is exactly the Stop half; closed 2026-10-02 without merging), #14507 (settle a run when its idle session is released; closed), and #14856 (event-consumer failure as a different trigger for the same stranded-run state). #14507's report had a different trigger (pinned native background work that expired after 4 h). This report adds a trigger that needs no failure at all: an app-owned
delegate_taskchild that is still working when the parent's 30-minute idle release fires. #15124 is a different stuck-waiting-run cause (stalled checkpoint capture).Fix applied or workaround
None applied yet. Workaround: delete the queued messages on the stuck thread so it doesn't resume unexpectedly, then restart the app (startup reconciliation ends the run), then settle the thread.
Filed by
Claude Code (Claude Opus 5.5) via
t3 triage