Repository navigation
[Bug]: "Waiting on subagent" chip never clears after the SDK reports the subagent stopped, and Stop does nothing #17847
Description
Activity
Note
Grok responding on behalf of Julius.
Thanks for the detailed report. I read through
main(read-only). Here's what I found. Line numbers are approximate.stoppedis treated as terminal. Inapps/server/src/orchestration-v2/Adapters/ClaudeAdapterV2.ts(~L3408),claudeTaskOutcomemaps SDKstoppedtocancelled. So your first guess, thatstoppedisn't treated as terminal, doesn't look like the cause.The notification probably can't find the node.
opaqueTaskWakeReport(~L3411) looks up thetask_idonly in in-memory per-process maps (pendingBackgroundTasksByNativeThread/lastKnownOpaqueTasks). After a restart, or a handoff to another Claude provider instance, those maps start empty. The task then resolves toundefinedand is reported as a genericcommandwith nochildThreadId. If that's right, nothing settles the persisted subagent node, so the projection keeps itrunningand the chip stays. That fits your second guess, that the node is orphaned by the session or instance change. I haven't reproduced it, though.Why Stop is a no-op.
ProviderTurnControlService.interruptroutes a subagent Stop to the adapter'sstopSubagent(~L8101). That function looks the task up in the in-memorysessionSubagentsByTaskIdand doesreturnwhen it isn't there. That's a silent success: no error, and nothing settles the projected node. It errors only when the subagent is known but has no live owning query. So for a subagent from an earlier process or instance, Stop seems to do nothing, by design. Unlike the turn-level Stop, which has a follow-up that settles the projection, nothing repairs the stale subagent row afterwards.Likely fix direction (unverified): fall back to the persisted node when a
task_notificationarrives for an unknowntask_id, and/or have Stop settle a subagent node in the projection ascancelledwhen no live query owns it. Another option is to sweep subagent nodes whose owning provider session isstoppedorerror.Related: #16355 (same stale-Running symptom after
completed), #17154, #14802, #12694, closed #7585. Open PRs #16852 and #16022 may touch nearby code. I haven't checked whether either covers the unknown-task or orphaned-node path.- addedvia-triageFiled through npx t3 triageFiled through npx t3 triagebugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.
on Oct 10, 2026 Same symptom on T3 Code (Nightly)
0.0.46-nightly.20261010.2908(macOS, Claude provider, Opus 5.5). Two details differ from the original report:- No server restart and no instance handoff. The T3 server has been up since 2026-10-09 19:59 MDT (
server-runtime.jsonstartedAt01:59:30Z). Only the thread's Claude session was replaced: a newclaudeprocess started 2026-10-10 09:28 MDT on the same provider instance. So the server-side maps weren't reset by a restart here, if they live for the server's whole lifetime. - The orphan is a Workflow-tool task, not an
Agentcall. The timeline itemtask:wi53hrhb5:subagent(typesubagent, text "Prepare: prepare") still readsrunning. It was last updated 2026-10-10T02:37:42Z, and its run (ordinal:16) iscompleted.
On the new session's first turn, the SDK delivered
<status>stopped</status>forwi53hrhb5("didn't finish before the previous session ended"). The item stayedrunning, and the sidebar still shows it as a subagent running for 12+ hours. No process for it exists.TaskStop("wi53hrhb5")in the new session returnsNo task found with ID: wi53hrhb5, so the agent can't send a stop either. A new workflow run in the same thread started normally beneath it.- No server restart and no instance handoff. The T3 server has been up since 2026-10-09 19:59 MDT (
Timeline (UTC), to rule out the provider-instance handoff as the cause:
- 2026-10-05 17:41: background subagent launched (Claude instance A).
- 2026-10-06 00:22: last write to the subagent's transcript; the session process ended with it unfinished.
- 2026-10-06 13:28: same instance A resumes and the SDK delivers
task-notificationwithstatus: stopped. - 2026-10-06 17:30: thread handed off to Claude instance B.
So the stale chip already existed while the thread was still on instance A. The
stoppednotification arrived on the original instance about 4 hours before the handoff and still didn't clear it.
Before submitting
completed), [Bug]: Cancelling a delegated task leaves its Claude native subagents running forever (task stuck waiting_for_children) #17154 (cancelling a delegated task leaves native subagents running).Area
apps/server (orchestrator subagent state) + apps/web (composer "Waiting on subagent" chip)
Steps to reproduce
Agentwithrun_in_background), e.g. titled "Ship 5790 small fixes then front page".task-notificationfor that agent with<status>stopped</status>and summary "Background agent … didn't finish before the previous session ended" / "No completion record was found for it in the previous session".Expected behavior
stoppedtask notification (or a subagent whose owning session is gone) clears the subagent's running state, and the composer chip disappears.Actual behavior
psshows nothing).Impact
Cosmetic but confusing: it looks like work is still running, and the user can't dismiss it. The chip also implies the thread is busy, which matters for anyone deciding whether a thread is safe to archive.
Version or commit
T3 Code (Nightly) 0.0.46-nightly.20261004.2652
Environment
macOS 26 (Darwin 25.6.0), Claude Agent provider (Opus 5.5), worktree-bound thread, two Claude provider instances (the thread was handed off between them).
Logs or stack traces
From the session transcript (SDK side), the notification T3 received on resume:
My guess:
stopped(as opposed tocompleted/failed) isn't treated as terminal for the subagent node, or the node belongs to the old provider instance after the handoff, so neither the notification nor Stop reaches it.Screenshots, recordings, or supporting files
Workaround
None found. Stop doesn't clear it; starting new turns doesn't clear it.