What happened
On a headless Linux host running T3 Code as the background service, the server process crashed twice and was restarted by systemd. Every running agent turn was cancelled. Both crashes happened after the auto-update from 0.0.46-nightly.20261008.2819 to 0.0.46-nightly.20261010.2948 (installed 2026-10-11 00:55 CEST). The first came at 01:13:33 and the second at 09:25:31 CEST, about 8 hours apart. The user never saw this crash on earlier builds.
Workload at the time:
- 6–12 long-running Claude threads (
claude-opus-5-5, full-access runtime), each in its own git worktree.
- A scheduled dispatcher thread that launches threads through the T3 MCP tools (
t3_thread_launch etc.).
- Clients connected through the T3 Connect managed tunnel (app.t3.codes web and the mobile app).
After each restart, every thread that had been running shows as cancelled and is not resumed. V2 orchestration recovery completed logged terminalizedRuns: 5, stoppedSessions: 7 after the first crash and terminalizedRuns: 10, stoppedSessions: 13 after the second.
Diagnosis
The crash. Both crashes end with the same uncaught exception. An 'error' event is emitted on a ClientHttp2Stream that has no listener, so Node throws it and exits with code 1. The service launcher then logs Active child exited unexpectedly (1). The error originates in the undici copy bundled with Node (node:internal/deps/undici), meaning fetch, not the server's own HTTP stack:
onHttp2SessionIdleTimeout (bundled undici 8.10.2) destroys an idle HTTP/2 session's socket with InformationalError("socket idle timeout").
- Any
ClientHttp2Stream still attached to that session gets destroyed with the error. If that stream has no 'error' listener at that moment, Node turns it into an uncaught exception.
- In this bundled undici,
buildConnector defaults allowH2 to true. Every HTTPS fetch to an h2-capable origin therefore goes over HTTP/2, without the app opting in. A grep of apps/ and packages/ at the tag finds no allowH2 or setGlobalDispatcher.
The runtime did not change between builds. Both binaries report node v26.8.2 and undici 8.10.2, checked with process.versions via a --require preload. They also ship the same native node_modules set. So something in the app between v0.0.46-nightly.20261008.2819 and v0.0.46-nightly.20261010.2948 (222 commits) started producing the HTTP/2 stream pattern that hits this path.
Most likely trigger. The outbound HTTPS traffic visible in the trace before the 09:25 crash is mostly GitHubApi.send → POST /graphql (about 28 calls in the preceding 2 minutes, all status 200). It also includes AgentAwarenessRelay/relay fetches and provider version lookups. GitHub is h2-capable, so it is the prime candidate. That code moved to @t3tools/source-control-github in #17607 with only cosmetic changes, so the move alone doesn't explain the regression. The one cross-cutting HTTP-layer change in the range is the Effect 4.0.1 → 4.0.2 upgrade (#17571). It touches HttpClient / HttpClientResponse (span status on 4xx/5xx, response handling) but not FetchHttpClient.ts. I could not pin the exact commit. The last ~19 s of trace before each crash were not flushed to disk, and the 01:13 trace has already rotated away.
Why it is fatal. The server has no process-level guard (process.on("uncaughtException") / uncaughtExceptionMonitor) in apps/server/src or packages/shared/src. One stray HTTP/2 stream error from a background fetch therefore kills the whole server and every agent session with it. After the restart, interrupted turns are terminalized rather than resumed. #16794 (ECONNRESET on a provider pipe) is the same class of failure.
Context, possibly unrelated:
- Before the first crash there was
event loop stalled for 2966 ms.
- From 00:59 to 01:07 there were four
rejected MCP request with an unusable credential { reason: 'unknown_or_expired_token' } warnings. The agents' MCP calls were apparently still carrying credentials issued before the 00:55 update restart.
- The second crash came right after a burst of
environment api request failed (errorTag: 'Interrupt') and cloudflared Incoming request ended abruptly: context canceled warnings. Those were inbound relay requests from a client being cancelled.
- No kernel OOM occurred. The systemd unit's memory peak at the 09:25 exit was 2.8G.
Suggested fixes:
- Stop the trigger. Pass an explicit undici dispatcher with
allowH2: false to the server's fetch (or FetchHttpClient.Fetch), or otherwise make sure every HTTP/2 stream keeps an 'error' listener until it closes. Also consider reporting the idle-timeout / listener gap to nodejs/undici.
- Defense in depth: one background HTTP request should not be able to take down every agent. Consider a narrowly scoped
uncaughtException handler that logs and survives UND_ERR_INFO / socket-level errors.
- After an unexpected restart, offer to resume or retry turns that recovery terminalized, instead of leaving them cancelled.
Steps to reproduce
No deterministic repro yet. It is intermittent: 2 crashes in about 8 hours.
- Run
0.0.46-nightly.20261010.2948 as the background service on Linux (bundled Node v26.8.2).
- Connect via T3 Connect. Keep 6–12 Claude threads running in worktrees of a GitHub-hosted repo, so the GitHub GraphQL / PR status polling is active, plus a dispatcher thread that launches threads through the T3 MCP tools.
- Wait. Eventually the server exits with the
socket idle timeout / UND_ERR_INFO stack below and all running threads become cancelled.
A targeted repro might be: issue fetches to an h2 origin (for example api.github.com/graphql), interrupt or abandon some responses mid-body, then let the session sit idle past the keep-alive timeout.
Version
0.0.46-nightly.20261010.2948 (regression vs 0.0.46-nightly.20261008.2819)
Environment
Ubuntu 24.04.5 LTS, Linux 6.8 x64; bundled Node v26.8.2 (undici 8.10.2); systemd user service (Restart=always); Claude Code CLI 2.1.293; T3 Connect managed tunnel
Evidence
# boot-service.log, crash 1 (01:13:33 CEST)
[00:59:39.301] WARN: rejected MCP request with an unusable credential { reason: 'unknown_or_expired_token' }
[01:00:08.262] WARN: rejected MCP request with an unusable credential { reason: 'unknown_or_expired_token' }
[01:02:30.343] WARN: rejected MCP request with an unusable credential { reason: 'unknown_or_expired_token' }
[01:07:37.265] WARN: rejected MCP request with an unusable credential { reason: 'unknown_or_expired_token' }
[01:09:06.451] WARN: event loop stalled for 2966 ms
node:events:505
throw er; // Unhandled 'error' event
^
InformationalError: socket idle timeout
at Timeout.onHttp2SessionIdleTimeout [as _onTimeout] (node:internal/deps/undici/undici:8822:19)
at listOnTimeout (node:internal/timers:687:11)
at process.processTimers (node:internal/timers:618:7)
Emitted 'error' event on ClientHttp2Stream instance at:
at emitErrorNT (node:internal/streams/destroy:170:8)
at emitErrorCloseNT (node:internal/streams/destroy:129:3)
at process.processTicksAndRejections (node:internal/process/task_queues:90:21) {
code: 'UND_ERR_INFO'
}
Node.js v26.8.2
[service-launcher] Active child exited unexpectedly (1).
[01:13:41.171] INFO: V2 orchestration recovery completed { terminalizedRuns: 5, stoppedSessions: 7, ... }
# crash 2 (09:25:31 CEST): preceded by relay request cancellations, then the identical stack
[09:25:23.468] WARN: Relay client reported a transport warning (cloudflared: "Incoming request ended abruptly: context canceled" for /api/orchestration/shell and /api/auth/session)
[09:25:23.504] WARN: environment api request failed { endpoint: 'HttpApiEndpoint', errorTag: 'Interrupt', ... }
node:events:505
throw er; // Unhandled 'error' event
InformationalError: socket idle timeout
at Timeout.onHttp2SessionIdleTimeout [as _onTimeout] (node:internal/deps/undici/undici:8822:19)
Emitted 'error' event on ClientHttp2Stream instance ... { code: 'UND_ERR_INFO' }
Node.js v26.8.2
[service-launcher] Active child exited unexpectedly (1).
[09:25:40.933] INFO: V2 orchestration recovery completed { terminalizedRuns: 10, stoppedSessions: 13, ... }
# journalctl --user -u t3code.service
Oct 11 01:13:33 t3code.service: Main process exited, code=exited, status=1/FAILURE
Oct 11 09:25:31 t3code.service: Main process exited, code=exited, status=1/FAILURE
Oct 11 09:25:31 t3code.service: Consumed 37min 1.222s CPU time, 2.8G memory peak, 35.0M memory swap peak.
# Embedded runtime of both builds (process.versions via --require preload)
0.0.46-nightly.20261008.2819: {"node":"v26.8.2","undici":"8.10.2"}
0.0.46-nightly.20261010.2948: {"node":"v26.8.2","undici":"8.10.2"}
# Node v26.8.2 deps/undici/undici.js: HTTP/2 is on by default for fetch
allowH2 = allowH2 != null ? allowH2 : true; // buildConnector
const err = new InformationalError("socket idle timeout"); // onHttp2SessionIdleTimeout
socket[kError] = err;
util.destroy(socket, err);
# server.trace.ndjson, 120 s before crash 2 (outbound HTTP spans)
28 x GitHubApi.send POST /graphql -> 200 (last at ~-60 s; trace's final ~19 s were not flushed)
197 x AgentAwarenessRelay.publishThread
1 x resolveLatestProviderVersion, 1 x readOpenCodeGoUsageLimits (~-19.5 s)
Related issues
#16794: headless t3 serve crashes on an unhandled read ECONNRESET and stops all 44 sessions. That is the same "one uncaught stream error kills every agent" failure mode, but a different source (provider child pipe vs undici HTTP/2 fetch). #17846 (closed) is similar. #5621 (closed) was an earlier unhandled-socket-error crash loop caused by an Effect regression. No existing issue mentions socket idle timeout, UND_ERR_INFO, or HTTP/2 fetch crashes. Nothing in v0.0.46-nightly.20261011.2955 or current main touches this path.
Fix applied or workaround
None applied. Possible workaround: roll back to 0.0.46-nightly.20261008.2819, which ran the same Node/undici without crashing for this user.
Filed by
Claude Code (claude-opus-5-5) via t3 triage
What happened
On a headless Linux host running T3 Code as the background service, the server process crashed twice and was restarted by systemd. Every running agent turn was cancelled. Both crashes happened after the auto-update from
0.0.46-nightly.20261008.2819to0.0.46-nightly.20261010.2948(installed 2026-10-11 00:55 CEST). The first came at 01:13:33 and the second at 09:25:31 CEST, about 8 hours apart. The user never saw this crash on earlier builds.Workload at the time:
claude-opus-5-5, full-access runtime), each in its own git worktree.t3_thread_launchetc.).After each restart, every thread that had been running shows as cancelled and is not resumed.
V2 orchestration recovery completedloggedterminalizedRuns: 5, stoppedSessions: 7after the first crash andterminalizedRuns: 10, stoppedSessions: 13after the second.Diagnosis
The crash. Both crashes end with the same uncaught exception. An
'error'event is emitted on aClientHttp2Streamthat has no listener, so Node throws it and exits with code 1. The service launcher then logsActive child exited unexpectedly (1). The error originates in the undici copy bundled with Node (node:internal/deps/undici), meaningfetch, not the server's own HTTP stack:onHttp2SessionIdleTimeout(bundled undici 8.10.2) destroys an idle HTTP/2 session's socket withInformationalError("socket idle timeout").ClientHttp2Streamstill attached to that session gets destroyed with the error. If that stream has no'error'listener at that moment, Node turns it into an uncaught exception.buildConnectordefaultsallowH2totrue. Every HTTPSfetchto an h2-capable origin therefore goes over HTTP/2, without the app opting in. A grep ofapps/andpackages/at the tag finds noallowH2orsetGlobalDispatcher.The runtime did not change between builds. Both binaries report
node v26.8.2andundici 8.10.2, checked withprocess.versionsvia a--requirepreload. They also ship the same nativenode_modulesset. So something in the app betweenv0.0.46-nightly.20261008.2819andv0.0.46-nightly.20261010.2948(222 commits) started producing the HTTP/2 stream pattern that hits this path.Most likely trigger. The outbound HTTPS traffic visible in the trace before the 09:25 crash is mostly
GitHubApi.send→POST /graphql(about 28 calls in the preceding 2 minutes, all status 200). It also includesAgentAwarenessRelay/relay fetches and provider version lookups. GitHub is h2-capable, so it is the prime candidate. That code moved to@t3tools/source-control-githubin #17607 with only cosmetic changes, so the move alone doesn't explain the regression. The one cross-cutting HTTP-layer change in the range is the Effect 4.0.1 → 4.0.2 upgrade (#17571). It touchesHttpClient/HttpClientResponse(span status on 4xx/5xx, response handling) but notFetchHttpClient.ts. I could not pin the exact commit. The last ~19 s of trace before each crash were not flushed to disk, and the 01:13 trace has already rotated away.Why it is fatal. The server has no process-level guard (
process.on("uncaughtException")/uncaughtExceptionMonitor) inapps/server/srcorpackages/shared/src. One stray HTTP/2 stream error from a background fetch therefore kills the whole server and every agent session with it. After the restart, interrupted turns are terminalized rather than resumed. #16794 (ECONNRESET on a provider pipe) is the same class of failure.Context, possibly unrelated:
event loop stalled for 2966 ms.rejected MCP request with an unusable credential { reason: 'unknown_or_expired_token' }warnings. The agents' MCP calls were apparently still carrying credentials issued before the 00:55 update restart.environment api request failed(errorTag: 'Interrupt') and cloudflaredIncoming request ended abruptly: context canceledwarnings. Those were inbound relay requests from a client being cancelled.Suggested fixes:
allowH2: falseto the server'sfetch(orFetchHttpClient.Fetch), or otherwise make sure every HTTP/2 stream keeps an'error'listener until it closes. Also consider reporting the idle-timeout / listener gap to nodejs/undici.uncaughtExceptionhandler that logs and survivesUND_ERR_INFO/ socket-level errors.Steps to reproduce
No deterministic repro yet. It is intermittent: 2 crashes in about 8 hours.
0.0.46-nightly.20261010.2948as the background service on Linux (bundled Node v26.8.2).socket idle timeout/UND_ERR_INFOstack below and all running threads become cancelled.A targeted repro might be: issue
fetches to an h2 origin (for exampleapi.github.com/graphql), interrupt or abandon some responses mid-body, then let the session sit idle past the keep-alive timeout.Version
0.0.46-nightly.20261010.2948 (regression vs 0.0.46-nightly.20261008.2819)
Environment
Ubuntu 24.04.5 LTS, Linux 6.8 x64; bundled Node v26.8.2 (undici 8.10.2); systemd user service (
Restart=always); Claude Code CLI 2.1.293; T3 Connect managed tunnelEvidence
Related issues
#16794: headless
t3 servecrashes on an unhandledread ECONNRESETand stops all 44 sessions. That is the same "one uncaught stream error kills every agent" failure mode, but a different source (provider child pipe vs undici HTTP/2 fetch). #17846 (closed) is similar. #5621 (closed) was an earlier unhandled-socket-error crash loop caused by an Effect regression. No existing issue mentionssocket idle timeout,UND_ERR_INFO, or HTTP/2 fetch crashes. Nothing inv0.0.46-nightly.20261011.2955or currentmaintouches this path.Fix applied or workaround
None applied. Possible workaround: roll back to
0.0.46-nightly.20261008.2819, which ran the same Node/undici without crashing for this user.Filed by
Claude Code (claude-opus-5-5) via
t3 triage