Skip to content

Server exits on unhandled 'socket idle timeout' (UND_ERR_INFO) from HTTP/2 fetch since 0.0.46-nightly.20261010.2948, cancelling all running threads #18141

Description

@szokeptr

What happened

On a headless Linux host running T3 Code as the background service, the server process crashed twice and was restarted by systemd. Every running agent turn was cancelled. Both crashes happened after the auto-update from 0.0.46-nightly.20261008.2819 to 0.0.46-nightly.20261010.2948 (installed 2026-10-11 00:55 CEST). The first came at 01:13:33 and the second at 09:25:31 CEST, about 8 hours apart. The user never saw this crash on earlier builds.

Workload at the time:

  • 6–12 long-running Claude threads (claude-opus-5-5, full-access runtime), each in its own git worktree.
  • A scheduled dispatcher thread that launches threads through the T3 MCP tools (t3_thread_launch etc.).
  • Clients connected through the T3 Connect managed tunnel (app.t3.codes web and the mobile app).

After each restart, every thread that had been running shows as cancelled and is not resumed. V2 orchestration recovery completed logged terminalizedRuns: 5, stoppedSessions: 7 after the first crash and terminalizedRuns: 10, stoppedSessions: 13 after the second.

Diagnosis

The crash. Both crashes end with the same uncaught exception. An 'error' event is emitted on a ClientHttp2Stream that has no listener, so Node throws it and exits with code 1. The service launcher then logs Active child exited unexpectedly (1). The error originates in the undici copy bundled with Node (node:internal/deps/undici), meaning fetch, not the server's own HTTP stack:

  • onHttp2SessionIdleTimeout (bundled undici 8.10.2) destroys an idle HTTP/2 session's socket with InformationalError("socket idle timeout").
  • Any ClientHttp2Stream still attached to that session gets destroyed with the error. If that stream has no 'error' listener at that moment, Node turns it into an uncaught exception.
  • In this bundled undici, buildConnector defaults allowH2 to true. Every HTTPS fetch to an h2-capable origin therefore goes over HTTP/2, without the app opting in. A grep of apps/ and packages/ at the tag finds no allowH2 or setGlobalDispatcher.

The runtime did not change between builds. Both binaries report node v26.8.2 and undici 8.10.2, checked with process.versions via a --require preload. They also ship the same native node_modules set. So something in the app between v0.0.46-nightly.20261008.2819 and v0.0.46-nightly.20261010.2948 (222 commits) started producing the HTTP/2 stream pattern that hits this path.

Most likely trigger. The outbound HTTPS traffic visible in the trace before the 09:25 crash is mostly GitHubApi.send → POST /graphql (about 28 calls in the preceding 2 minutes, all status 200). It also includes AgentAwarenessRelay/relay fetches and provider version lookups. GitHub is h2-capable, so it is the prime candidate. That code moved to @t3tools/source-control-github in #17607 with only cosmetic changes, so the move alone doesn't explain the regression. The one cross-cutting HTTP-layer change in the range is the Effect 4.0.1 → 4.0.2 upgrade (#17571). It touches HttpClient / HttpClientResponse (span status on 4xx/5xx, response handling) but not FetchHttpClient.ts. I could not pin the exact commit. The last ~19 s of trace before each crash were not flushed to disk, and the 01:13 trace has already rotated away.

Why it is fatal. The server has no process-level guard (process.on("uncaughtException") / uncaughtExceptionMonitor) in apps/server/src or packages/shared/src. One stray HTTP/2 stream error from a background fetch therefore kills the whole server and every agent session with it. After the restart, interrupted turns are terminalized rather than resumed. #16794 (ECONNRESET on a provider pipe) is the same class of failure.

Context, possibly unrelated:

  • Before the first crash there was event loop stalled for 2966 ms.
  • From 00:59 to 01:07 there were four rejected MCP request with an unusable credential { reason: 'unknown_or_expired_token' } warnings. The agents' MCP calls were apparently still carrying credentials issued before the 00:55 update restart.
  • The second crash came right after a burst of environment api request failed (errorTag: 'Interrupt') and cloudflared Incoming request ended abruptly: context canceled warnings. Those were inbound relay requests from a client being cancelled.
  • No kernel OOM occurred. The systemd unit's memory peak at the 09:25 exit was 2.8G.

Suggested fixes:

  1. Stop the trigger. Pass an explicit undici dispatcher with allowH2: false to the server's fetch (or FetchHttpClient.Fetch), or otherwise make sure every HTTP/2 stream keeps an 'error' listener until it closes. Also consider reporting the idle-timeout / listener gap to nodejs/undici.
  2. Defense in depth: one background HTTP request should not be able to take down every agent. Consider a narrowly scoped uncaughtException handler that logs and survives UND_ERR_INFO / socket-level errors.
  3. After an unexpected restart, offer to resume or retry turns that recovery terminalized, instead of leaving them cancelled.

Steps to reproduce

No deterministic repro yet. It is intermittent: 2 crashes in about 8 hours.

  1. Run 0.0.46-nightly.20261010.2948 as the background service on Linux (bundled Node v26.8.2).
  2. Connect via T3 Connect. Keep 6–12 Claude threads running in worktrees of a GitHub-hosted repo, so the GitHub GraphQL / PR status polling is active, plus a dispatcher thread that launches threads through the T3 MCP tools.
  3. Wait. Eventually the server exits with the socket idle timeout / UND_ERR_INFO stack below and all running threads become cancelled.

A targeted repro might be: issue fetches to an h2 origin (for example api.github.com/graphql), interrupt or abandon some responses mid-body, then let the session sit idle past the keep-alive timeout.

Version

0.0.46-nightly.20261010.2948 (regression vs 0.0.46-nightly.20261008.2819)

Environment

Ubuntu 24.04.5 LTS, Linux 6.8 x64; bundled Node v26.8.2 (undici 8.10.2); systemd user service (Restart=always); Claude Code CLI 2.1.293; T3 Connect managed tunnel

Evidence

# boot-service.log, crash 1 (01:13:33 CEST)
[00:59:39.301] WARN: rejected MCP request with an unusable credential { reason: 'unknown_or_expired_token' }
[01:00:08.262] WARN: rejected MCP request with an unusable credential { reason: 'unknown_or_expired_token' }
[01:02:30.343] WARN: rejected MCP request with an unusable credential { reason: 'unknown_or_expired_token' }
[01:07:37.265] WARN: rejected MCP request with an unusable credential { reason: 'unknown_or_expired_token' }
[01:09:06.451] WARN: event loop stalled for 2966 ms
node:events:505
    throw er; // Unhandled 'error' event
    ^

InformationalError: socket idle timeout
    at Timeout.onHttp2SessionIdleTimeout [as _onTimeout] (node:internal/deps/undici/undici:8822:19)
    at listOnTimeout (node:internal/timers:687:11)
    at process.processTimers (node:internal/timers:618:7)
Emitted 'error' event on ClientHttp2Stream instance at:
    at emitErrorNT (node:internal/streams/destroy:170:8)
    at emitErrorCloseNT (node:internal/streams/destroy:129:3)
    at process.processTicksAndRejections (node:internal/process/task_queues:90:21) {
  code: 'UND_ERR_INFO'
}

Node.js v26.8.2
[service-launcher] Active child exited unexpectedly (1).
[01:13:41.171] INFO: V2 orchestration recovery completed { terminalizedRuns: 5, stoppedSessions: 7, ... }

# crash 2 (09:25:31 CEST): preceded by relay request cancellations, then the identical stack
[09:25:23.468] WARN: Relay client reported a transport warning  (cloudflared: "Incoming request ended abruptly: context canceled" for /api/orchestration/shell and /api/auth/session)
[09:25:23.504] WARN: environment api request failed { endpoint: 'HttpApiEndpoint', errorTag: 'Interrupt', ... }
node:events:505
    throw er; // Unhandled 'error' event
InformationalError: socket idle timeout
    at Timeout.onHttp2SessionIdleTimeout [as _onTimeout] (node:internal/deps/undici/undici:8822:19)
Emitted 'error' event on ClientHttp2Stream instance ... { code: 'UND_ERR_INFO' }
Node.js v26.8.2
[service-launcher] Active child exited unexpectedly (1).
[09:25:40.933] INFO: V2 orchestration recovery completed { terminalizedRuns: 10, stoppedSessions: 13, ... }

# journalctl --user -u t3code.service
Oct 11 01:13:33 t3code.service: Main process exited, code=exited, status=1/FAILURE
Oct 11 09:25:31 t3code.service: Main process exited, code=exited, status=1/FAILURE
Oct 11 09:25:31 t3code.service: Consumed 37min 1.222s CPU time, 2.8G memory peak, 35.0M memory swap peak.

# Embedded runtime of both builds (process.versions via --require preload)
0.0.46-nightly.20261008.2819: {"node":"v26.8.2","undici":"8.10.2"}
0.0.46-nightly.20261010.2948: {"node":"v26.8.2","undici":"8.10.2"}

# Node v26.8.2 deps/undici/undici.js: HTTP/2 is on by default for fetch
allowH2 = allowH2 != null ? allowH2 : true;                          // buildConnector
const err = new InformationalError("socket idle timeout");            // onHttp2SessionIdleTimeout
socket[kError] = err;
util.destroy(socket, err);

# server.trace.ndjson, 120 s before crash 2 (outbound HTTP spans)
28 x GitHubApi.send  POST /graphql -> 200   (last at ~-60 s; trace's final ~19 s were not flushed)
197 x AgentAwarenessRelay.publishThread
1 x resolveLatestProviderVersion, 1 x readOpenCodeGoUsageLimits (~-19.5 s)

Related issues

#16794: headless t3 serve crashes on an unhandled read ECONNRESET and stops all 44 sessions. That is the same "one uncaught stream error kills every agent" failure mode, but a different source (provider child pipe vs undici HTTP/2 fetch). #17846 (closed) is similar. #5621 (closed) was an earlier unhandled-socket-error crash loop caused by an Effect regression. No existing issue mentions socket idle timeout, UND_ERR_INFO, or HTTP/2 fetch crashes. Nothing in v0.0.46-nightly.20261011.2955 or current main touches this path.

Fix applied or workaround

None applied. Possible workaround: roll back to 0.0.46-nightly.20261008.2819, which ran the same Node/undici without crashing for this user.

Filed by

Claude Code (claude-opus-5-5) via t3 triage

Activity

  1. juliusmarminge commented on Oct 11, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Confirmed bug on current main: an unhandled 'error' on an undici ClientHttp2Stream (InformationalError: socket idle timeout / UND_ERR_INFO) kills the whole server process, and recovery terminalizes every running turn.

    What main does. Outbound HTTPS goes through Effect FetchHttpClient.layer → Node's bundled fetch/undici. There is no allowH2: false / setGlobalDispatcher anywhere under apps/ or packages/. There is also no process.on("uncaughtException") / uncaughtExceptionMonitor in apps/server or packages/shared. The only existing mitigation of this failure class is inbound-only: httpResponseErrorGuard.ts attaches 'error' listeners to HTTP server responses and upgrade sockets so a client reset cannot take the process down. Outbound fetch streams are uncovered.

    Upstream. Same stack is tracked as nodejs/undici#5936. Node 26 defaults allowH2 to true, so every HTTPS fetch to an h2 origin (including api.github.com) is on this path.

    Not a duplicate of #16794. That issue (and open #16555) is an unhandled 'error' on a provider child pipe (read ECONNRESET). Same blast radius, different emitter. Fixing the child-pipe listeners will not stop this undici idle-timeout crash. Cross-link them for a shared process-level guard if we add one.

    Regression trigger: hypothesis only. #17607 moved GitHubApi.send into @t3tools/source-control-github but still uses Effect HttpClient the same way. #17571 (Effect 4.0.1 → 4.0.2) is the one cross-cutting HTTP-layer bump in the reported window; I did not pin a commit that newly leaves a stream without an 'error' listener. Nothing between 20261010.2948 and 20261011.2955 touches this path.

    Useful next steps (no promise/timeline): disable H2 on the server's fetch dispatcher, and/or add a narrowly scoped process guard for UND_ERR_INFO / socket-level uncaught errors so one background request cannot cancel every agent.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 11, 2026
  3. diyaavirmani commented on Oct 11, 2026

    @diyaavirmani

    Hi @juliusmarminge , I’d like to work on this issue. I’ve read CONTRIBUTING.md and AGENTS.md, along with the report and your triage response.

    I’d start by building a focused reproduction of the HTTP/2 idle-timeout crash, then investigate disabling HTTP/2 for the server’s outbound fetch dispatcher and add regression coverage. I’d keep changes to process-level exception handling and thread recovery outside the initial PR.

    If nobody is already working on this, could you please assign it to me and confirm whether that scope works for you? Thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions