Skip to content

[Bug]: A resumed Claude subagent's child thread stays empty when a later run resumes it #16261

Description

@Artem-B

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. In a Claude thread, send:

    Call the Agent tool once: subagent_type "general-purpose", model "haiku", run_in_background true, description "Repro sub C", prompt: "Make 4 Bash tool calls strictly one after another: put exactly one tool call in each message and wait for its result before the next. Each call runs exactly: python3 -c 'import time; time.sleep(5)'. After the fourth result, reply with exactly: READY_C". After the call returns, reply with exactly LAUNCHED_C and stop.

  2. When LAUNCHED_C appears, and while the subagent still runs, send:
    1. Call Bash in the foreground: python3 -c 'import time; time.sleep(35)'. 2) Use SendMessage to send "Repro sub C" (use its agent id): "Make 5 Bash tool calls strictly one after another, one tool call per message. Each call runs exactly: python3 -c 'import time; time.sleep(6)'. Then reply with exactly: DONE_C". 3) Call Bash: python3 -c 'import time; time.sleep(15)'. 4) Reply with exactly SENT_C and stop.
  3. Wait for the subagent to finish and for the parent to report it.
  4. Open the subagent's child thread.

The subagent must complete during the run of step 2. That run then resumes it. The 35 s sleep makes sure of this.

Expected behavior

After the resume, the child thread shows the SendMessage prompt, the 5 Bash calls and DONE_C. They appear live while a parent turn is active, or at the latest when the subagent completes.

Actual behavior

The child thread stops at READY_C, the end of the first run. It never gets the resume prompt, the Bash calls or DONE_C, also after completion. The subagent card in the parent thread does update, and shows completed with result DONE_C.

The items are not stored, so a reload does not show them. In real use this looks like a hung subagent: an implementer that I resumed with a decision did 40 minutes of work, and its thread showed nothing after its earlier report.

Cause, as far as I can tell from the source at 7812230:

  • routeProviderEvent (RunExecutionService.ts L349) keeps a child-thread turn_item.updated only if the thread is in the run's ownedThreadIds.
  • A run owns a child thread when it sees the app_thread.created for it (it launched the subagent, L376). Or, canRouteRelatedSubagent (L179, used in ProviderTurnStartService.ts L944) passes it at run start, but only for a completed subagent.
  • The sequence: run A launches the subagent. Run B starts while the subagent is running, so B does not take the thread. The subagent completes during B, and A's ingestion fiber stops. B resumes the subagent with SendMessage.
  • On the resume, ClaudeAdapterV2 sets the subagent's runId to run B (ClaudeAdapterV2.ts L4111). This keeps B's fiber alive and routes the card. But the subagent.updated case (RunExecutionService.ts L419) does not add childThreadId to ownedThreadIds. After this, no fiber owns the child thread.
  • Later runs start while the subagent is running, so they do not take it either.

A possible fix: when routeProviderEvent accepts a subagent.updated because the run owns event.subagent.runId, add event.subagent.childThreadId to ownedThreadIds. Check the case where the launch run's fiber is still alive for other background work, because both fibers would then own the thread.

The existing claude_background_subagent_lifecycle fixture resumes a subagent from a run that started after its completion, so it does not catch this.

Impact

Major degradation or frequent failure

Version or commit

0.0.46-nightly.20261005.2676 (7812230)

Environment

Linux server (systemd user service), web client. Claude Code 2.1.289. Parent model claude-opus-5-5 (original case) and claude-sonnet-5-5 (repro). Subagent model haiku (repro) and inherited from the parent (original case).

Logs or stack traces

# provider/events.<thread>.log, Sub C (UTC). parent = root frames, SUB = parent_tool_use_id set
23:36:47.085 task_started   task=a142ab40880aa733f tool_use=<Agent call>      # run 11 launches
23:36:48.650 result                                                            # run 11 ends
23:36:53.886 init                                                              # run 12 starts, subagent running
23:37:14.811 task_notification task=a142ab40880aa733f completed                # completes inside run 12
23:37:33.247 task_started   task=a142ab40880aa733f tool_use=<SendMessage call> # run 12 resumes it
23:37:34.894 SUB assistant tool_use Bash      ... 5 Bash calls ...
23:37:49.376 result                                                            # run 12 ends
23:38:10.828 SUB assistant text DONE_C
23:38:10.858 task_notification task=a142ab40880aa733f completed
23:38:12.305 result                                                            # wake run 13 ends

# orchestration_v2_projection_turn_items for the child thread
1  user_message      23:36:47.085  "Make 4 Bash tool calls ..."
2-10 reasoning / command_execution  23:36:48 - 23:37:13
11 assistant_message 23:37:14.795  READY_C
(nothing after this)

# orchestration_v2_projection_subagents: status completed, runId ...:ordinal:12, result "DONE_C"


Two control cases on the same thread work. If the same run launches and resumes the subagent, the child thread gets every item. If the subagent completes while the parent is idle and a later run resumes it, the later run takes the thread at start, and the items appear.

Screenshots, recordings, or supporting files

No response

Workaround

Use delegate_task for long subagent work, or read the subagent's transcript JSONL under ~/.claude/projects/.../subagents/. Do not resume a native subagent from a run that started before the subagent's last completion. This is hard to control, because the parent usually answers the completion report in the same turn.

Activity

  1. added
    bugSomething is broken or behaving incorrectly.
    needs-triageIssue needs maintainer review and initial categorization.
    on Oct 5, 2026
  2. juliusmarminge commented on Oct 5, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks @Artem-B for the detailed repro and the run-by-run sequence. It made this much easier to trace.

    What I found

    On current main (c138163c8c), the child thread's items are dropped before they're stored, so a reload can't bring them back.

    • Child-thread items are written with runId: null (makeSubagentConversationArtifacts, and Claude tool calls under a subagent in ensureToolCallStarted).
    • routeProviderEvent keeps a turn_item.updated only when the run owns that runId or the item's thread is already in ownedThreadIds.
    • A run gains the child thread in only two places: app_thread.created (when it launches the subagent), or relatedThreadIds at run start, where canRouteRelatedSubagent is true only for completed subagents.

    That lines up with your sequence. Run 11 launches the subagent and owns the child thread. Run 12 starts while it's still running, so it doesn't take the thread. The subagent completes during run 12, and the SendMessage reopen reattributes subagent.runId to run 12 (updateClaudeSubagentNode). The parent card keeps updating because subagent.updated is accepted via ownsRun, but nothing adds childThreadId to run 12's ownedThreadIds. Run 13 then starts while the subagent is running again, so it doesn't take the thread either. The resume prompt, the Bash calls, and DONE_C are all dropped.

    The existing claude_background_subagent_lifecycle test doesn't cover this overlap, since its resume runs only after the subagent completed and the wake run settled. This also looks distinct from #13668, which handled a resume run starting after the subagent was already completed.

    Likely fix area

    One option is enrolling childThreadId into ownedThreadIds when routeProviderEvent accepts a subagent.updated via ownsRun(event.subagent.runId). Two things to watch with that approach:

    1. In updateClaudeSubagentNode, the child root node.updated and the SendMessage prompt are emitted before subagent.updated, so enrolling only on that event would still drop the prompt unless the reattributed subagent.updated is routed first.
    2. If the launch fiber is still alive for other background work, it still owns the child thread. It would need to release the thread when runId moves, or the same events could be stored twice (the reason canRouteRelatedSubagent exists).

    Your two control cases (the same run launching and resuming the subagent, and a later run resuming only after completion while the parent is idle) are worth keeping as regression checks.

    A maintainer will decide on the fix direction.

  3. added
    via-triageFiled through npx t3 triage
    and removed
    needs-triageIssue needs maintainer review and initial categorization.
    on Oct 5, 2026
  4. eimexdev commented on Oct 7, 2026

    @eimexdev
    Contributor

    Another way to hit this, with event-log evidence from a real session: stopping a Claude subagent and then resuming it with SendMessage. Besides the empty child thread, this also leaves the child thread's footer reading "Cancelled" while the parent's Agents list shows it Working.

    Sequence (from orchestration_events)

    Run What happened
    17 starts The subagent is completed, so run 17 takes over its child thread.
    17 SendMessage resumes it. The child root turn goes to running and the prompt reaches the child thread.
    18 starts The subagent is still running, so run 18 doesn't take over the child thread.
    (17's fiber) The parent stops the subagent. The parent's subagent node and the child root turn both go to cancelled, and run 17's fiber ends.
    18 SendMessage resumes it. updateClaudeSubagentNode moves the subagent to run 18 and emits the parent's subagent node, the child root node.updated, the resume prompt and subagent.updated. Only the parent-side events are stored: sequence 89294 (parent node running) is followed directly by 89295 (subagent.updated), with nothing for the child thread in between. On the earlier resume in run 17, the child root update and prompt were stored (89173/89174).

    Visible result

    • Parent Agents list: Working, with fresh progress. It reads the parent-side subagent record.
    • Child thread: the composer bar reads "Cancelled · Runs on its own". deriveProviderSubagentStatus reads the child's runless root turn, which never got the running update. The thread stops at the previous prompt. This happens on web and mobile.

    Two points for the fix

    1. Stopped subagents fail even without the run overlap. canRouteRelatedSubagent only passes completed, and its comment says "An interrupted, failed or cancelled one is never resumed". Claude's SendMessage does resume them. So stop, then a new run, then SendMessage, loses the child thread even when the new run starts after the stop.
    2. The ordering caveat from the triage applies here. The child root node.updated and the prompt are emitted before the moved subagent.updated. Taking over the child thread only when that event arrives would still drop both, and the bar would still say "Cancelled".
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions