Repository navigation
[Bug]: Subagent status stays red after a usage limit recovery, and never updates while running #7314
Description
Activity
Confirmed on the current published nightly, with a different teardown path: an unexpected T3 server
SIGKILLwhile a Codex subagent was running.Environment
- T3 Code
0.0.39-nightly.20260906.1303 - Linux
7.0.14-14-pvex86_64, web client in Chromium 152 - Node
v24.20.0 - Codex provider
Reproduction
- Start T3 against a persistent T3 home and begin a Codex turn.
- Spawn a subagent and let it publish
task.updated { status: "running" }followed bytask.progressevents. SIGKILLthe captured T3 server process while that subagent is in a tool call, then restart T3 against the same home.- Resume the parent thread and start another subagent.
- Compare the Agents panel, the thread activity summary, and the composer status.
Persisted evidence
Thread
461e3498-1e1e-4cf7-9ed1-c2b5a5186db9had parent turn01a077d3-7d7a-7ca0-b230-e326064240a5and child task01a077d7-45b0-7463-bd42-c09ca8074d80(vue_core_parity). Times below are UTC:18:36:51.897 task.updated status=running 18:45:38.958 task.progress /bin/bash -lc 'npm run knip' 18:45:39 T3 process killed by the OOM killer (SIGKILL) 18:45:46.180 parent projection_turns state=error, completed_at setThere is no terminal
task.updated/task.completedevent for the child after the crash. On replay, the parent turn is correctly terminal, but the child task remains folded as working from its pre-crash events even though its native process cannot have survived.At 18:51 UTC, after restart, the same web screen simultaneously showed:
- Agents panel:
vue_core_parityand the newdocs_reconcilechild both had blue working indicators. - Thread activity summary:
1 working. - Composer status:
2 agents working. - Native subagent registry: only
docs_reconcilewas running.
This looks like two parts of the same defect described in this issue:
- session/server teardown does not synthesize a terminal state for live child tasks; and
- the panel/composer and thread liveness summary derive working state from different folds, so they disagree after replay.
Expected: when a parent turn is finalized as
errorduring restart recovery, any non-terminal child tasks owned by that turn should also become terminal (interrupted/stopped), and all client surfaces should derive the working count from the same reconciled state.This is related to the broader session reconciliation cases in #4713, but here
projection_turnsfor the parent is already correct; the stale state is specifically the child-task liveness represented by persisted task activities.- T3 Code
Confirmed on T3 Code (Alpha) 0.0.40, macOS 15.7.3, Claude provider (claude-fable-5-1).
Same sequence as Part 1, with one difference in how the agents came back. After the limit reset, the parent resumed each failed subagent explicitly (SendMessage to the same agent ID). They were not automatically re-dispatched. Three subagents, same result for all.
From
events.<threadId>.log, one of them:15:40:54 task.started taskId=a9548a234e20953d1 15:47:03 task.updated status=failed (session limit, HTTP 429) 15:47:03 task.completed status=failed 16:01:06 task.started taskId=a9548a234e20953d1 <- same taskId, resumed 16:09:58 task.updated status=completed 16:09:58 task.completed status=completedOne extra detail: the row's token and tool counts did update to the final values of the resumed run. Status, colour and the error text stayed frozen from the first failure. So the panel is consuming the later events for usage but not for status.
Confirmed on my setup on 2026-09-29. I hit the five-hour usage limit while several subagents were running. After usage became available again, I sent “Continue” in the chat. The agents restarted, but the Agents panel still showed the old red session-limit error for some of them while other agents showed completed or active states. The stale error text and status remained visible after the agents had resumed, matching Part 1 of this issue. The attached screenshot shows the mixed state: agents B, C, and F still show the API/session-limit error, while other entries show completed or active-looking statuses.
kvnloo commented
on Sep 29, 2026 ContributorMore actionsPart 1 now has a pretty crisp canonical mechanism in #14002. Claude resume does not necessarily emit
task.updated(status=running); it emits anothertask.startedfor the sametaskId, with a newtoolUseIdfrom theSendMessageactivation. The client fold currently treats post-terminal starts as delayed metadata, which explains the stale red/completed row while usage keeps updating.Draft #14031 tests the smallest discriminator I could find: remember observed start
toolUseIds per task and reopen a terminal task only whentask.startedcarries an unseen start identity. A repeated/late start with an already-seentoolUseIdremains terminal, so the existing out-of-order protection stays intact.I would treat that as the fix for this issue's Part 1 and keep #7314 open for Part 2 (live current-step/progress updates), which is a separate projection problem. That avoids closing this broader thread just because resumed-status recovery lands.
Part 1 has the same cause as #14002. Claude resumes the subagent with another
task_startedunder theSendMessagecall'stool_use_id, and the Agents panel fold treated it as late metadata. #14696 fixes this in the fold, with a client-runtime-only change. The resumed row now reopens as running, clears the old error, and settles on its own outcome. The fold doesn't read the error text, so a session-limit failure behaves the same as the network failures I tested with. In a replay of my own persisted rows, 8 resumed subagents went fromfailed, run 1, to run 2 with no stale error.#14696 doesn't touch parts 2 and 3.
Before submitting
Area
apps/server,apps/web,packages/client-runtimeSummary
Subagent status in the Agents panel is unreliable in two ways:
Both come from data the app already receives and then drops.
Note on existing issues
The session-level half of this is already known: #6513, with PRs #7165, #5077 and #5473 in flight (and #6639 closed). None of them touch subagent status — they all work at the thread/session layer, and none change
subagentRuntime.tsor the Agents panel. Part 1 below is the subagent-level gap those fixes leave open. #7128 is the opposite direction (progress wrongly reviving an idle task), so please read part 2 alongside it.Part 1 — a subagent that recovers from a usage limit stays red
Steps to reproduce
Actual behavior
The subagents are marked
failed, and they stay red for the rest of the session — including while they are working again and after they complete.From my own logs (
~/.t3/userdata/logs/provider/events.<threadId>.log), four subagents died on one limit and all four came back:Expected behavior
A usage limit is a pause, not a failure. The row should show something like "waiting · resets 2:20pm", and it must go back to normal once the agent resumes.
Why it happens
Three things line up:
The rate-limit signal is discarded.
ClaudeAdapter.ts:3474andCodexAdapter.ts:1401both emitaccount.rate-limits.updatedcarrying the window type andresetsAt. Grepping the tree for that string returns only the two emitters and the contract —ProviderRuntimeIngestion.tshas no case arm, so it falls through. The UI can't know a limit happened. (PR fix(server): type the account rate-limit runtime payload #5473 makes the same observation.)It collapses into
failed. With no rate-limit check in the LLM path, the result hits the defaultreturn "failed"atClaudeAdapter.ts:956. There's also no better value available: the status union atsubagentRuntime.ts:22has no throttled/waiting-for-quota member.failedis frozen, and re-dispatch reuses the taskId.subagentRuntime.ts:515blocks a progress tick from reopening a terminal agent:Only an explicit
running/pendingstatus event reopens the row, so the recovered agent inherits the old terminal state under the same id.Blast radius: one stuck member turns the whole collapsed workflow dot red (
AgentsPanel.tsx:482) and the chat CTA dot too (MessagesTimeline.tsx:2171).For contrast, the codebase already models this properly for the GitHub CLI —
packages/contracts/src/vcs.ts:76has a first-class"rate-limited"kind. The LLM path never got it.Part 2 — a running subagent gets no status updates
Actual behavior
Every canonical event for one subagent across a 7-minute run:
Status is set at spawn, then nothing for 7m19s. A later run in the same log has a 30m12s gap.
In that same 7-minute window the SDK sent 128
task_progressnotifications, each carrying the current step:{"task_id":"a76927b010c77a27d","description":"Running Inspect worktree layout", "last_tool_name":"Bash","usage":{"tool_uses":1,"duration_ms":3419}}Across all my logs: 1309 progress notifications, 65 status updates. The progress events are projected into token usage only —
descriptionandlast_tool_nameare dropped.Expected behavior
The panel reflects the agent's most recent real activity. The data needed for this is already arriving many times per minute.
Also contributing
ProviderRuntimeIngestion.ts:592writes under a stable idtask-progress:{threadId}:{taskId}, so one row exists per task ever. If an event is missed, the last value is pinned with nothing to notice it's old.AgentsPanel.tsx:101updatestextContentonly and never re-folds).stopSessionInternal(ClaudeAdapter.ts:3595) doesn't drainliveTaskIds, so an agent that ends without atask_notificationshows running forever.Part 3 — small related bug: retryable Codex errors paint the red pill
CodexSessionRuntime.ts:1428:The
lastErrorspread sits outside thewillRetryternary, so a retryable error leaves the sessionrunningwithlastErrorset. The sidebar red pill keys purely offlastErrorbeing non-null (Sidebar.tsx:337), andProviderRuntimeIngestion.ts:1592only clearslastErrorwhenstatus === "ready"—turn.startedmaps to"running", so resuming never clears it.Claude's equivalent path does the right thing:
ClaudeAdapter.ts:3325turnsapi_retryintosession.state.changed{state:"running"}to keep the session visibly alive.Suggestions
account.rate-limits.updatedinto a non-terminal throttled state at the subagent level, not just the thread banner the open PRs add.failedagent (subagentRuntime.ts:515) so red can clear itself. Please check this against [Bug]: Finished threads stay Working after delayed task progress #7128 — that issue wants progress to not revive an idle task, so the two need a shared rule.task_progressinto status, not only token usage. The current step is already on the wire ~18x/minute.lastActivityAtso a quiet agent looks different from a busy one.Items 1 and 3 need no new plumbing — only reading events the app already emits.
Impact
Moderate: the Agents panel can't be trusted, so I wait on work that's already done, or kill agents that were fine.
Version or commit
0.0.34-nightly.20260814.1092Environment
macOS (Darwin 25.5.0) desktop app, Claude and Codex providers.
How I traced this
The nightly ships sourcemaps with full
sourcesContent, so all line numbers above refer to real source, not minified offsets. You may want to check whether that's intended for released builds.