Skip to content

[Bug]: Backend crash leaves Claude agents running after their runs are marked cancelled #15357

Description

@allisonmahmood

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Summary

When the desktop backend dies abruptly (SIGKILL), the claude processes it spawned keep running. On restart, recovery marks their runs cancelled and assumes the processes are gone. They aren't. The agents carry on calling tools: running git, editing files, posting to GitHub. None of their output reaches T3. The user sees a cancelled thread while the agent keeps changing their machine.

Visual explainer with the full timeline: https://post.patchyhq.com/d/9yc2qhd8o4yy

Steps to reproduce

  1. In a Claude thread, ask the agent to run until [ -e /tmp/release ]; do sleep 1; done in the foreground, then run touch /tmp/acted and reply "done".
  2. While it waits, kill -9 the backend from a terminal outside T3 (its pid is in ~/.t3/userdata/server-runtime.json). Let the desktop restart it.
  3. Wait for the run to show as cancelled, then touch /tmp/release.

Expected behavior

The agent process dies with the backend, or recovery kills it before marking the run cancelled.

Actual behavior

The agent process is reparented to the user's systemd manager and keeps going. It creates /tmp/acted and replies "done". T3 shows neither.

From our repro (UTC):

Time T3 recorded The agent did
21:50:57.092 Backend killed. The agent shares the backend's pgid and sid. Waiting in its shell loop
21:50:58.861 Run cancelled, provider thread idle, no assistant item Still alive, ppid now 1167 (systemd --user)
21:51:24.508 Nothing Next tool call: writes /tmp/acted
21:51:25.488 Nothing Final reply written to its Claude transcript

In the original incident, two mid-turn agents went on for 4 and 7 minutes after the crash and filed GitHub issues that never appeared in T3. An idle agent with a Claude Code Monitor running also survived. Its Monitor fired 91 seconds after the crash, and the orphan ran a whole new turn on its own for about 10 minutes. Another idle agent, waiting on a plain background command, exited instead (that case is #15358). So what an orphan does depends on Claude Code internals.

Orphans reach the restarted backend but get refused. Their MCP client reconnects and gets 401 invalid_mcp_credential, which is right. The problem is that they're still alive.

Cause

  • ClaudeAdapterV2.ts:614 opens Claude through the SDK's query(). The SDK spawns the CLI without detached, so it shares the backend's process group and session, and nothing kills it when the backend dies abruptly. Graceful shutdown closes the query in a finalizer (ClaudeAdapterV2.ts:7314), but SIGKILL skips that.
  • ProviderRuntimeRecoveryService.ts:570 says // All provider processes are gone on startup/shutdown, then marks the runs cancelled (:306). There is no pid ledger, no reaping of orphans, and no way to reattach.

#14926 hit the same behaviour in the test suite: the fake Claude CLI was "reparented to PID 1 and runs until reboot", so it now exits when stdin closes. The real CLI keeps working after stdin closes mid-turn. PR #15323 handles SIGINT and SIGTERM but not SIGKILL, so orphans aren't covered there either.

Possible fix

Make that comment true. Either set PR_SET_PDEATHSIG on provider children (or spawn them through a small wrapper that does), or record each provider's pid and start time and kill survivors during startup recovery, before marking runs cancelled.

Impact

Major degradation or frequent failure

Version or commit

0.0.46-nightly.20261003.2632 (f391794), Orchestrator V2

Environment

Linux 7.2.5 (Arch-based), desktop AppImage, Claude Code 2.1.288 via Agent SDK 0.3.276, Claude Opus 5.5

Logs or stack traces

# Claude Code MCP log, 11 seconds after the backend was killed
Error POSTing to endpoint: {"error":"invalid_mcp_credential","message":"A valid provider-scoped MCP bearer credential is required."}
Server rejected the configured Authorization header (HTTP 401).

# desktop.trace.ndjson, runBackendProcess exit cause
Process interrupted due to receipt of signal: 'SIGKILL'

Workaround

None in T3. After a backend crash, check for claude processes whose parent is pid 1 (or your systemd user manager) and kill them by hand.

Related reports from the same incident: #15358 (a thread waiting on background work stays stranded after a restart), #15359 (a killed backend is logged as "Failed to read the exit status").

Filed by Allison Mahmood. Investigated and drafted with Claude Opus 5.5 in Claude Code.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions