Skip to content

[Bug]: Subagent status stays red after a usage limit recovery, and never updates while running #7314

Description

@ChiChuRita

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server, apps/web, packages/client-runtime

Summary

Subagent status in the Agents panel is unreliable in two ways:

  1. After a usage limit, a subagent goes red and stays red even after it recovers and finishes the work.
  2. While a subagent runs, its status never updates — the current step is set at spawn and not touched again until the run ends.

Both come from data the app already receives and then drops.

Note on existing issues

The session-level half of this is already known: #6513, with PRs #7165, #5077 and #5473 in flight (and #6639 closed). None of them touch subagent status — they all work at the thread/session layer, and none change subagentRuntime.ts or the Agents panel. Part 1 below is the subagent-level gap those fixes leave open. #7128 is the opposite direction (progress wrongly reviving an idle task), so please read part 2 alongside it.


Part 1 — a subagent that recovers from a usage limit stays red

Steps to reproduce

  1. Spawn several subagents on a Claude session.
  2. Let the account hit its 5-hour usage limit mid-run.
  3. Wait for the limit window to reset and the subagents to be re-dispatched.
  4. Watch the Agents panel.

Actual behavior

The subagents are marked failed, and they stay red for the rest of the session — including while they are working again and after they complete.

From my own logs (~/.t3/userdata/logs/provider/events.<threadId>.log), four subagents died on one limit and all four came back:

09:38:56.716  account.rate-limits.updated  {"status":"rejected","rateLimitType":"five_hour","resetsAt":1786969200}
09:38:56.718  task.updated  taskId=aa7ad7db11045589e  status=failed
              error="Agent terminated early due to an API error: You've hit your session limit · resets 2:20pm"
09:49:23.309  task.started  taskId=aa7ad7db11045589e   <- same taskId, working again
10:00:45.545  task.completed  status=completed         <- finished; row still red

Expected behavior

A usage limit is a pause, not a failure. The row should show something like "waiting · resets 2:20pm", and it must go back to normal once the agent resumes.

Why it happens

Three things line up:

  • The rate-limit signal is discarded. ClaudeAdapter.ts:3474 and CodexAdapter.ts:1401 both emit account.rate-limits.updated carrying the window type and resetsAt. Grepping the tree for that string returns only the two emitters and the contract — ProviderRuntimeIngestion.ts has no case arm, so it falls through. The UI can't know a limit happened. (PR fix(server): type the account rate-limit runtime payload #5473 makes the same observation.)

  • It collapses into failed. With no rate-limit check in the LLM path, the result hits the default return "failed" at ClaudeAdapter.ts:956. There's also no better value available: the status union at subagentRuntime.ts:22 has no throttled/waiting-for-quota member.

  • failed is frozen, and re-dispatch reuses the taskId. subagentRuntime.ts:515 blocks a progress tick from reopening a terminal agent:

    } else if (
      (payload.usageSnapshot !== true || !existed) &&
      !isTerminalSubagentStatus(agent.status) &&   // <- blocks recovery
      agent.status !== "idle"
    ) {
      applyStatus(agent, "running", at);
    }

    Only an explicit running/pending status event reopens the row, so the recovered agent inherits the old terminal state under the same id.

Blast radius: one stuck member turns the whole collapsed workflow dot red (AgentsPanel.tsx:482) and the chat CTA dot too (MessagesTimeline.tsx:2171).

For contrast, the codebase already models this properly for the GitHub CLI — packages/contracts/src/vcs.ts:76 has a first-class "rate-limited" kind. The LLM path never got it.


Part 2 — a running subagent gets no status updates

Actual behavior

Every canonical event for one subagent across a 7-minute run:

09:32:26  task.started
09:37:40  thread.token-usage.updated   (x4)
09:39:45  task.updated  status=failed

Status is set at spawn, then nothing for 7m19s. A later run in the same log has a 30m12s gap.

In that same 7-minute window the SDK sent 128 task_progress notifications, each carrying the current step:

{"task_id":"a76927b010c77a27d","description":"Running Inspect worktree layout",
 "last_tool_name":"Bash","usage":{"tool_uses":1,"duration_ms":3419}}

Across all my logs: 1309 progress notifications, 65 status updates. The progress events are projected into token usage only — description and last_tool_name are dropped.

Expected behavior

The panel reflects the agent's most recent real activity. The data needed for this is already arriving many times per minute.

Also contributing

  • Progress rows are collapse-on-write. ProviderRuntimeIngestion.ts:592 writes under a stable id task-progress:{threadId}:{taskId}, so one row exists per task ever. If an event is missed, the last value is pinned with nothing to notice it's old.
  • No liveness timer anywhere, client or server. So a finished agent reads "Working" with a happily ticking elapsed clock (AgentsPanel.tsx:101 updates textContent only and never re-folds).
  • Terminal events aren't synthesized on teardown. stopSessionInternal (ClaudeAdapter.ts:3595) doesn't drain liveTaskIds, so an agent that ends without a task_notification shows running forever.

Part 3 — small related bug: retryable Codex errors paint the red pill

CodexSessionRuntime.ts:1428:

return updateSession(sessionRef, {
  status: willRetry ? "running" : "error",
  ...(errorMessage ? { lastError: errorMessage } : {}),
});

The lastError spread sits outside the willRetry ternary, so a retryable error leaves the session running with lastError set. The sidebar red pill keys purely off lastError being non-null (Sidebar.tsx:337), and ProviderRuntimeIngestion.ts:1592 only clears lastError when status === "ready" — turn.started maps to "running", so resuming never clears it.

Claude's equivalent path does the right thing: ClaudeAdapter.ts:3325 turns api_retry into session.state.changed{state:"running"} to keep the session visibly alive.


Suggestions

  1. Consume account.rate-limits.updated into a non-terminal throttled state at the subagent level, not just the thread banner the open PRs add.
  2. Let a live progress tick reopen a failed agent (subagentRuntime.ts:515) so red can clear itself. Please check this against [Bug]: Finished threads stay Working after delayed task progress #7128 — that issue wants progress to not revive an idle task, so the two need a shared rule.
  3. Project task_progress into status, not only token usage. The current step is already on the wire ~18x/minute.
  4. Add a lastActivityAt so a quiet agent looks different from a busy one.

Items 1 and 3 need no new plumbing — only reading events the app already emits.

Impact

Moderate: the Agents panel can't be trusted, so I wait on work that's already done, or kill agents that were fine.

Version or commit

0.0.34-nightly.20260814.1092

Environment

macOS (Darwin 25.5.0) desktop app, Claude and Codex providers.

How I traced this

The nightly ships sourcemaps with full sourcesContent, so all line numbers above refer to real source, not minified offsets. You may want to check whether that's intended for released builds.

Activity

  1. Dropje97 commented on Sep 6, 2026

    @Dropje97

    Confirmed on the current published nightly, with a different teardown path: an unexpected T3 server SIGKILL while a Codex subagent was running.

    Environment

    • T3 Code 0.0.39-nightly.20260906.1303
    • Linux 7.0.14-14-pve x86_64, web client in Chromium 152
    • Node v24.20.0
    • Codex provider

    Reproduction

    1. Start T3 against a persistent T3 home and begin a Codex turn.
    2. Spawn a subagent and let it publish task.updated { status: "running" } followed by task.progress events.
    3. SIGKILL the captured T3 server process while that subagent is in a tool call, then restart T3 against the same home.
    4. Resume the parent thread and start another subagent.
    5. Compare the Agents panel, the thread activity summary, and the composer status.

    Persisted evidence

    Thread 461e3498-1e1e-4cf7-9ed1-c2b5a5186db9 had parent turn 01a077d3-7d7a-7ca0-b230-e326064240a5 and child task 01a077d7-45b0-7463-bd42-c09ca8074d80 (vue_core_parity). Times below are UTC:

    18:36:51.897  task.updated   status=running
    18:45:38.958  task.progress  /bin/bash -lc 'npm run knip'
    18:45:39      T3 process killed by the OOM killer (SIGKILL)
    18:45:46.180  parent projection_turns state=error, completed_at set
    

    There is no terminal task.updated/task.completed event for the child after the crash. On replay, the parent turn is correctly terminal, but the child task remains folded as working from its pre-crash events even though its native process cannot have survived.

    At 18:51 UTC, after restart, the same web screen simultaneously showed:

    • Agents panel: vue_core_parity and the new docs_reconcile child both had blue working indicators.
    • Thread activity summary: 1 working.
    • Composer status: 2 agents working.
    • Native subagent registry: only docs_reconcile was running.

    This looks like two parts of the same defect described in this issue:

    1. session/server teardown does not synthesize a terminal state for live child tasks; and
    2. the panel/composer and thread liveness summary derive working state from different folds, so they disagree after replay.

    Expected: when a parent turn is finalized as error during restart recovery, any non-terminal child tasks owned by that turn should also become terminal (interrupted/stopped), and all client surfaces should derive the working count from the same reconciled state.

    This is related to the broader session reconciliation cases in #4713, but here projection_turns for the parent is already correct; the stale state is specifically the child-task liveness represented by persisted task activities.

  2. michaelcummings commented on Sep 21, 2026

    @michaelcummings

    Confirmed on T3 Code (Alpha) 0.0.40, macOS 15.7.3, Claude provider (claude-fable-5-1).

    Same sequence as Part 1, with one difference in how the agents came back. After the limit reset, the parent resumed each failed subagent explicitly (SendMessage to the same agent ID). They were not automatically re-dispatched. Three subagents, same result for all.

    From events.<threadId>.log, one of them:

    15:40:54  task.started    taskId=a9548a234e20953d1
    15:47:03  task.updated    status=failed   (session limit, HTTP 429)
    15:47:03  task.completed  status=failed
    16:01:06  task.started    taskId=a9548a234e20953d1   <- same taskId, resumed
    16:09:58  task.updated    status=completed
    16:09:58  task.completed  status=completed
    

    One extra detail: the row's token and tool counts did update to the final values of the resumed run. Status, colour and the error text stayed frozen from the first failure. So the panel is consuming the later events for usage but not for status.

  3. nabertronic commented on Sep 29, 2026

    @nabertronic

    Confirmed on my setup on 2026-09-29. I hit the five-hour usage limit while several subagents were running. After usage became available again, I sent “Continue” in the chat. The agents restarted, but the Agents panel still showed the old red session-limit error for some of them while other agents showed completed or active states. The stale error text and status remained visible after the agents had resumed, matching Part 1 of this issue. The attached screenshot shows the mixed state: agents B, C, and F still show the API/session-limit error, while other entries show completed or active-looking statuses.

  4. kvnloo commented on Sep 29, 2026

    @kvnloo
    Contributor

    Part 1 now has a pretty crisp canonical mechanism in #14002. Claude resume does not necessarily emit task.updated(status=running); it emits another task.started for the same taskId, with a new toolUseId from the SendMessage activation. The client fold currently treats post-terminal starts as delayed metadata, which explains the stale red/completed row while usage keeps updating.

    Draft #14031 tests the smallest discriminator I could find: remember observed start toolUseIds per task and reopen a terminal task only when task.started carries an unseen start identity. A repeated/late start with an already-seen toolUseId remains terminal, so the existing out-of-order protection stays intact.

    I would treat that as the fix for this issue's Part 1 and keep #7314 open for Part 2 (live current-step/progress updates), which is a separate projection problem. That avoids closing this broader thread just because resumed-status recovery lands.

  5. navidnadali commented on Oct 2, 2026

    @navidnadali

    Part 1 has the same cause as #14002. Claude resumes the subagent with another task_started under the SendMessage call's tool_use_id, and the Agents panel fold treated it as late metadata. #14696 fixes this in the fold, with a client-runtime-only change. The resumed row now reopens as running, clears the old error, and settles on its own outcome. The fold doesn't read the error text, so a session-limit failure behaves the same as the network failures I tested with. In a replay of my own persisted rows, 8 resumed subagents went from failed, run 1, to run 2 with no stale error.

    #14696 doesn't touch parts 2 and 3.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions