Before submitting
Area
apps/server
Summary
When the desktop backend dies abruptly (SIGKILL), the claude processes it spawned keep running. On restart, recovery marks their runs cancelled and assumes the processes are gone. They aren't. The agents carry on calling tools: running git, editing files, posting to GitHub. None of their output reaches T3. The user sees a cancelled thread while the agent keeps changing their machine.
Visual explainer with the full timeline: https://post.patchyhq.com/d/9yc2qhd8o4yy
Steps to reproduce
- In a Claude thread, ask the agent to run
until [ -e /tmp/release ]; do sleep 1; done in the foreground, then run touch /tmp/acted and reply "done".
- While it waits,
kill -9 the backend from a terminal outside T3 (its pid is in ~/.t3/userdata/server-runtime.json). Let the desktop restart it.
- Wait for the run to show as cancelled, then
touch /tmp/release.
Expected behavior
The agent process dies with the backend, or recovery kills it before marking the run cancelled.
Actual behavior
The agent process is reparented to the user's systemd manager and keeps going. It creates /tmp/acted and replies "done". T3 shows neither.
From our repro (UTC):
| Time |
T3 recorded |
The agent did |
| 21:50:57.092 |
Backend killed. The agent shares the backend's pgid and sid. |
Waiting in its shell loop |
| 21:50:58.861 |
Run cancelled, provider thread idle, no assistant item |
Still alive, ppid now 1167 (systemd --user) |
| 21:51:24.508 |
Nothing |
Next tool call: writes /tmp/acted |
| 21:51:25.488 |
Nothing |
Final reply written to its Claude transcript |
In the original incident, two mid-turn agents went on for 4 and 7 minutes after the crash and filed GitHub issues that never appeared in T3. An idle agent with a Claude Code Monitor running also survived. Its Monitor fired 91 seconds after the crash, and the orphan ran a whole new turn on its own for about 10 minutes. Another idle agent, waiting on a plain background command, exited instead (that case is #15358). So what an orphan does depends on Claude Code internals.
Orphans reach the restarted backend but get refused. Their MCP client reconnects and gets 401 invalid_mcp_credential, which is right. The problem is that they're still alive.
Cause
ClaudeAdapterV2.ts:614 opens Claude through the SDK's query(). The SDK spawns the CLI without detached, so it shares the backend's process group and session, and nothing kills it when the backend dies abruptly. Graceful shutdown closes the query in a finalizer (ClaudeAdapterV2.ts:7314), but SIGKILL skips that.
ProviderRuntimeRecoveryService.ts:570 says // All provider processes are gone on startup/shutdown, then marks the runs cancelled (:306). There is no pid ledger, no reaping of orphans, and no way to reattach.
#14926 hit the same behaviour in the test suite: the fake Claude CLI was "reparented to PID 1 and runs until reboot", so it now exits when stdin closes. The real CLI keeps working after stdin closes mid-turn. PR #15323 handles SIGINT and SIGTERM but not SIGKILL, so orphans aren't covered there either.
Possible fix
Make that comment true. Either set PR_SET_PDEATHSIG on provider children (or spawn them through a small wrapper that does), or record each provider's pid and start time and kill survivors during startup recovery, before marking runs cancelled.
Impact
Major degradation or frequent failure
Version or commit
0.0.46-nightly.20261003.2632 (f391794), Orchestrator V2
Environment
Linux 7.2.5 (Arch-based), desktop AppImage, Claude Code 2.1.288 via Agent SDK 0.3.276, Claude Opus 5.5
Logs or stack traces
# Claude Code MCP log, 11 seconds after the backend was killed
Error POSTing to endpoint: {"error":"invalid_mcp_credential","message":"A valid provider-scoped MCP bearer credential is required."}
Server rejected the configured Authorization header (HTTP 401).
# desktop.trace.ndjson, runBackendProcess exit cause
Process interrupted due to receipt of signal: 'SIGKILL'
Workaround
None in T3. After a backend crash, check for claude processes whose parent is pid 1 (or your systemd user manager) and kill them by hand.
Related reports from the same incident: #15358 (a thread waiting on background work stays stranded after a restart), #15359 (a killed backend is logged as "Failed to read the exit status").
Filed by Allison Mahmood. Investigated and drafted with Claude Opus 5.5 in Claude Code.
Before submitting
Area
apps/server
Summary
When the desktop backend dies abruptly (SIGKILL), the
claudeprocesses it spawned keep running. On restart, recovery marks their runscancelledand assumes the processes are gone. They aren't. The agents carry on calling tools: running git, editing files, posting to GitHub. None of their output reaches T3. The user sees a cancelled thread while the agent keeps changing their machine.Visual explainer with the full timeline: https://post.patchyhq.com/d/9yc2qhd8o4yy
Steps to reproduce
until [ -e /tmp/release ]; do sleep 1; donein the foreground, then runtouch /tmp/actedand reply "done".kill -9the backend from a terminal outside T3 (its pid is in~/.t3/userdata/server-runtime.json). Let the desktop restart it.touch /tmp/release.Expected behavior
The agent process dies with the backend, or recovery kills it before marking the run cancelled.
Actual behavior
The agent process is reparented to the user's systemd manager and keeps going. It creates
/tmp/actedand replies "done". T3 shows neither.From our repro (UTC):
cancelled, provider threadidle, no assistant item/tmp/actedIn the original incident, two mid-turn agents went on for 4 and 7 minutes after the crash and filed GitHub issues that never appeared in T3. An idle agent with a Claude Code Monitor running also survived. Its Monitor fired 91 seconds after the crash, and the orphan ran a whole new turn on its own for about 10 minutes. Another idle agent, waiting on a plain background command, exited instead (that case is #15358). So what an orphan does depends on Claude Code internals.
Orphans reach the restarted backend but get refused. Their MCP client reconnects and gets
401 invalid_mcp_credential, which is right. The problem is that they're still alive.Cause
ClaudeAdapterV2.ts:614opens Claude through the SDK'squery(). The SDK spawns the CLI withoutdetached, so it shares the backend's process group and session, and nothing kills it when the backend dies abruptly. Graceful shutdown closes the query in a finalizer (ClaudeAdapterV2.ts:7314), but SIGKILL skips that.ProviderRuntimeRecoveryService.ts:570says// All provider processes are gone on startup/shutdown, then marks the runs cancelled (:306). There is no pid ledger, no reaping of orphans, and no way to reattach.#14926 hit the same behaviour in the test suite: the fake Claude CLI was "reparented to PID 1 and runs until reboot", so it now exits when stdin closes. The real CLI keeps working after stdin closes mid-turn. PR #15323 handles SIGINT and SIGTERM but not SIGKILL, so orphans aren't covered there either.
Possible fix
Make that comment true. Either set
PR_SET_PDEATHSIGon provider children (or spawn them through a small wrapper that does), or record each provider's pid and start time and kill survivors during startup recovery, before marking runs cancelled.Impact
Major degradation or frequent failure
Version or commit
0.0.46-nightly.20261003.2632 (f391794), Orchestrator V2
Environment
Linux 7.2.5 (Arch-based), desktop AppImage, Claude Code 2.1.288 via Agent SDK 0.3.276, Claude Opus 5.5
Logs or stack traces
Workaround
None in T3. After a backend crash, check for
claudeprocesses whose parent is pid 1 (or your systemd user manager) and kill them by hand.Related reports from the same incident: #15358 (a thread waiting on background work stays stranded after a restart), #15359 (a killed backend is logged as "Failed to read the exit status").
Filed by Allison Mahmood. Investigated and drafted with Claude Opus 5.5 in Claude Code.