You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
http: a send-failed request re-pools its dead connection — one WriteFailed poisons every later non-streaming request (compaction, [title], subagents) #177
At ~229,774 metered tokens auto-compaction failed with a full retry storm, and then — after /clear, on a fresh tiny conversation — the background title request failed the exact same way:
[compacting ~229774 tokens…]
[network error: WriteFailed — retrying in 250ms (1/6)]
… (2/6) … (6/6)
[request failed: WriteFailed — giving up this turn]
[auto-compaction failed: ApiError]
…
/clear
[title] [network error: WriteFailed — retrying in 250ms (1/6)] … (6/6)
A ~1MB compaction body and a one-message title body failing identically 6/6 in the same process is not a network flake.
Root cause: the non-streaming POST path can't poison a dead connection
Every non-live request — compaction (stream_quiet, agent_compact.zig:260), the [title] sub-agent (title.zig, .sub=true), all subagents — goes request() → postWatched → http.zig post() → std.http.Client.fetch on the shared root client (main.zig:268).
In Zig 0.16 std, a request whose body write fails leaves reader.state == .ready; Request.deinit maps that to closing = false (Client.zig:890-909) and release() returns the dead connection to the keep-alive pool; findConnection scans the free list LIFO (append :146, scan-from-last :91), so the same corpse is handed to every subsequent request to that host. client.fetch never exposes the Request, so post() has no way to mark it dirty.
We already diagnosed and fixed this exact livelock on the streaming path — agent_stream.zig:243-251:
// A failed SEND leaves reader.state == .ready, which Request.deinit// reads as "connection still clean" and returns it to the keep-alive// pool — every retry (and every later turn) then pulls the same dead// connection back out: HttpConnectionClosing once, WriteFailed// forever after (observed in the wild). Poison it on any error so// deinit discards it and the retry dials fresh.errdeferif (req.connection) |conn| { conn.closing=true; };
— but it was never mirrored into http.zig post(). Worse, the retry policy's doc comment (http.zig:197-198) claims "the connection is poisoned + re-dialed fresh on each retry (postStream/postWatched errdefer)" — true only for postStream. The retry loop was designed assuming a protection that the postWatched path never had.
So in the observed session: the oversized compaction body (that half is #174/#175's meter undercount — the full-input resend re-serializes all encrypted reasoning items and blew the ~272k wall, killed by the backend mid-write per the #163 comment at agent_compact.zig:250-254) failed once, its dead connection went back into the pool, and attempts 2-6 — plus the post-/clear title request, plus any subagent — kept pulling the same corpse. Live turns kept working because WS + postStreamFresh bypass the pool, which is exactly why only the quiet/background requests looked cursed.
Fix
Rewrite post() to use client.request(...) + the existing sendHeadTask helper instead of client.fetch, and add the same errdefer if (req.connection) |conn| conn.closing = true; poison (plus on the 429/5xx early-returns, mirroring postStream).
Correct the stale RetryPlan comment.
Follow-ups (separate): retrying a deterministically-rejected oversized body 6× with the identical bytes is futile — request() only rebuilds on CodexWsReanchor; once fix(codex): context meter tracks the full-history resend + recover from context-window rejections (#174) #175's meter fix lands the compaction body should stay under the wall, but WriteFailed-with-large-body deserves a shrink-and-rebuild rather than blind replay. Upstream candidate: std's Request.deinit could check the connection's latched writer error before re-pooling.
Observed (prod session, codex / gpt-5.6-sol, effort xhigh, v0.0.200)
At ~229,774 metered tokens auto-compaction failed with a full retry storm, and then — after
/clear, on a fresh tiny conversation — the background title request failed the exact same way:A ~1MB compaction body and a one-message title body failing identically 6/6 in the same process is not a network flake.
Root cause: the non-streaming POST path can't poison a dead connection
Every non-live request — compaction (
stream_quiet, agent_compact.zig:260), the[title]sub-agent (title.zig,.sub=true), all subagents — goesrequest()→postWatched→http.zig post()→std.http.Client.fetchon the shared root client (main.zig:268).In Zig 0.16 std, a request whose body write fails leaves
reader.state == .ready;Request.deinitmaps that toclosing = false(Client.zig:890-909) andrelease()returns the dead connection to the keep-alive pool;findConnectionscans the free list LIFO (append :146, scan-from-last :91), so the same corpse is handed to every subsequent request to that host.client.fetchnever exposes theRequest, sopost()has no way to mark it dirty.We already diagnosed and fixed this exact livelock on the streaming path — agent_stream.zig:243-251:
— but it was never mirrored into
http.zig post(). Worse, the retry policy's doc comment (http.zig:197-198) claims "the connection is poisoned + re-dialed fresh on each retry (postStream/postWatched errdefer)" — true only for postStream. The retry loop was designed assuming a protection that the postWatched path never had.So in the observed session: the oversized compaction body (that half is #174/#175's meter undercount — the full-input resend re-serializes all encrypted reasoning items and blew the ~272k wall, killed by the backend mid-write per the #163 comment at agent_compact.zig:250-254) failed once, its dead connection went back into the pool, and attempts 2-6 — plus the post-/clear title request, plus any subagent — kept pulling the same corpse. Live turns kept working because WS +
postStreamFreshbypass the pool, which is exactly why only the quiet/background requests looked cursed.Fix
post()to useclient.request(...)+ the existingsendHeadTaskhelper instead ofclient.fetch, and add the sameerrdefer if (req.connection) |conn| conn.closing = true;poison (plus on the 429/5xx early-returns, mirroring postStream).CodexWsReanchor; once fix(codex): context meter tracks the full-history resend + recover from context-window rejections (#174) #175's meter fix lands the compaction body should stay under the wall, but WriteFailed-with-large-body deserves a shrink-and-rebuild rather than blind replay. Upstream candidate: std'sRequest.deinitcould check the connection's latched writer error before re-pooling.Related: #163, #165, #174, #175.