Skip to content

[Bug]: "Waiting on subagent" chip never clears after the SDK reports the subagent stopped, and Stop does nothing #17847

Description

@jameswasher

Before submitting

Area

apps/server (orchestrator subagent state) + apps/web (composer "Waiting on subagent" chip)

Steps to reproduce

  1. In a Claude-provider thread, launch a depth-1 background subagent (Agent with run_in_background), e.g. titled "Ship 5790 small fixes then front page".
  2. The Claude session ends before the subagent finishes (here, the session hit API errors and was later replaced: the thread moved through compactions and a context handoff to another Claude provider instance / config dir).
  3. On the next session start the SDK delivers a task-notification for that agent with <status>stopped</status> and summary "Background agent … didn't finish before the previous session ended" / "No completion record was found for it in the previous session".
  4. Keep using the thread for days.

Expected behavior

  • A stopped task notification (or a subagent whose owning session is gone) clears the subagent's running state, and the composer chip disappears.
  • Failing that, the chip's Stop button clears it.

Actual behavior

  • The composer still shows "Waiting on subagent Ship 5790 small fixes then front page … Stop" five days later (launched 2026-10-05, last write to its transcript 2026-10-05 19:22 local).
  • No process for that agent or its session exists (ps shows nothing).
  • Clicking Stop does nothing; the chip stays.
  • The thread keeps working normally; new turns start and finish beneath the stale chip.

Impact

Cosmetic but confusing: it looks like work is still running, and the user can't dismiss it. The chip also implies the thread is busy, which matters for anyone deciding whether a thread is safe to archive.

Version or commit

T3 Code (Nightly) 0.0.46-nightly.20261004.2652

Environment

macOS 26 (Darwin 25.6.0), Claude Agent provider (Opus 5.5), worktree-bound thread, two Claude provider instances (the thread was handed off between them).

Logs or stack traces

From the session transcript (SDK side), the notification T3 received on resume:

<task-notification>
<task-id>ad0a7ffcd62d360aa</task-id>
<status>stopped</status>
<summary>Background agent "Ship 5790 small fixes then front page" didn't finish before the previous session ended</summary>
<note>No completion record was found for it in the previous session …</note>

My guess: stopped (as opposed to completed/failed) isn't treated as terminal for the subagent node, or the node belongs to the old provider instance after the handoff, so neither the notification nor Stop reaches it.

Screenshots, recordings, or supporting files

Stale "Waiting on subagent" chip with Stop

Workaround

None found. Stop doesn't clear it; starting new turns doesn't clear it.

Activity

  1. juliusmarminge commented on Oct 10, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Thanks for the detailed report. I read through main (read-only). Here's what I found. Line numbers are approximate.

    stopped is treated as terminal. In apps/server/src/orchestration-v2/Adapters/ClaudeAdapterV2.ts (~L3408), claudeTaskOutcome maps SDK stopped to cancelled. So your first guess, that stopped isn't treated as terminal, doesn't look like the cause.

    The notification probably can't find the node. opaqueTaskWakeReport (~L3411) looks up the task_id only in in-memory per-process maps (pendingBackgroundTasksByNativeThread / lastKnownOpaqueTasks). After a restart, or a handoff to another Claude provider instance, those maps start empty. The task then resolves to undefined and is reported as a generic command with no childThreadId. If that's right, nothing settles the persisted subagent node, so the projection keeps it running and the chip stays. That fits your second guess, that the node is orphaned by the session or instance change. I haven't reproduced it, though.

    Why Stop is a no-op. ProviderTurnControlService.interrupt routes a subagent Stop to the adapter's stopSubagent (~L8101). That function looks the task up in the in-memory sessionSubagentsByTaskId and does return when it isn't there. That's a silent success: no error, and nothing settles the projected node. It errors only when the subagent is known but has no live owning query. So for a subagent from an earlier process or instance, Stop seems to do nothing, by design. Unlike the turn-level Stop, which has a follow-up that settles the projection, nothing repairs the stale subagent row afterwards.

    Likely fix direction (unverified): fall back to the persisted node when a task_notification arrives for an unknown task_id, and/or have Stop settle a subagent node in the projection as cancelled when no live query owns it. Another option is to sweep subagent nodes whose owning provider session is stopped or error.

    Related: #16355 (same stale-Running symptom after completed), #17154, #14802, #12694, closed #7585. Open PRs #16852 and #16022 may touch nearby code. I haven't checked whether either covers the unknown-task or orphaned-node path.

  2. added
    via-triageFiled through npx t3 triage
    bugSomething is broken or behaving incorrectly.
    on Oct 10, 2026
  3. SamuelFranks commented on Oct 10, 2026

    @SamuelFranks

    Same symptom on T3 Code (Nightly) 0.0.46-nightly.20261010.2908 (macOS, Claude provider, Opus 5.5). Two details differ from the original report:

    • No server restart and no instance handoff. The T3 server has been up since 2026-10-09 19:59 MDT (server-runtime.json startedAt 01:59:30Z). Only the thread's Claude session was replaced: a new claude process started 2026-10-10 09:28 MDT on the same provider instance. So the server-side maps weren't reset by a restart here, if they live for the server's whole lifetime.
    • The orphan is a Workflow-tool task, not an Agent call. The timeline item task:wi53hrhb5:subagent (type subagent, text "Prepare: prepare") still reads running. It was last updated 2026-10-10T02:37:42Z, and its run (ordinal:16) is completed.

    On the new session's first turn, the SDK delivered <status>stopped</status> for wi53hrhb5 ("didn't finish before the previous session ended"). The item stayed running, and the sidebar still shows it as a subagent running for 12+ hours. No process for it exists. TaskStop("wi53hrhb5") in the new session returns No task found with ID: wi53hrhb5, so the agent can't send a stop either. A new workflow run in the same thread started normally beneath it.

  4. jameswasher commented on Oct 10, 2026

    @jameswasher
    Author

    Timeline (UTC), to rule out the provider-instance handoff as the cause:

    • 2026-10-05 17:41: background subagent launched (Claude instance A).
    • 2026-10-06 00:22: last write to the subagent's transcript; the session process ended with it unfinished.
    • 2026-10-06 13:28: same instance A resumes and the SDK delivers task-notification with status: stopped.
    • 2026-10-06 17:30: thread handed off to Claude instance B.

    So the stale chip already existed while the thread was still on instance A. The stopped notification arrived on the original instance about 4 hours before the handoff and still didn't clear it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions