fix(t3team): a dead agent turn can no longer settle the askAgent step as its answer - #284
Merged
johnnyelwailer merged 1 commit intoSep 16, 2026
Conversation
… as its answer A turn that died without completing (provider session exited mid-stream, or the turn was aborted by the host watchdog / provider) settled the step's ask with whatever the agent had streamed before the death — the truncated preamble was reported as the step's full answer, and the run continued. Only a session ending in "error" engaged the bounded re-drive → needs-attention path. - Turn tracker: any dead-state write that ends the wait (error / stopped / interrupted) now settles the ask as "failed" with a status-specific reason (or the write's lastError), so the reactor's bounded re-drive fires and, when the budget is spent, the run fails with the provider's reason. Only a ready/idle write (the turn actually completed) may settle with text. - findCompletedAnswer: messages belonging to the thread's latest turn no longer count as a completed answer when that turn ended interrupted/error — the re-drive cannot consume a dead turn's preamble. - ProviderRuntimeIngestion: removed the shadowed second `case "turn.aborted"` (→ "ready", unreachable dead code left by the 2026-09 merge); the live "interrupted" case documents why the status must not be "ready". The stale fork test pinning "ready" (failing on main since the merge) now pins "interrupted", matching upstream #9653's OpenCode tests. Regression tests: host-level integration tests pin that a mid-stream dead turn (stopped) and an aborted turn (interrupted) reject the askAgent step, re-drive it, and complete the run only with the re-driven turn's answer; tracker and lookup unit tests pin the per-status settlement.
johnnyelwailer
marked this pull request as ready for review
September 16, 2026 17:01
johnnyelwailer
deleted the
feature/host-failed-agent-turn-swallowed-as-step-completion-8429c93c
branch
September 16, 2026 17:01
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A failed agent turn silently completed the
askAgentstepProblem
When a workflow's
thread.askAgent(agent step) turn died without completing — the provider session exited mid-stream (session.exited), or the turn was aborted before it finished (host turn-watchdog, provider abort) — the step settled with whatever the agent had streamed up to that point. The truncated preamble was reported to the run as the step's full answer, and the run continued. The documented failure path (bounded re-drive → "needs attention" with the provider's reason) only engaged when the session ended in status"error".Reproduced live: a gateway stream break left the step's thread mid-turn; the step then "completed" on the one preamble message the agent had emitted before the break.
Root cause (three seams, all host-side)
"error"as a dead turn. Int3team-workflowTurnResolution.ts,noteSessionmarked the watch as failed only when the session write that ended the wait had status"error". The"stopped"(session exited) and"interrupted"(turn aborted) writes ended the wait but settled the ask as"answer"with the last streamed candidate — the partial text.findCompletedAnswer(used by the interrupted-turn re-drive and its invariant-retry path) returned the first completed assistant message after the prompt — which, for a dead turn, is its preamble. Even a correctly rejected turn could have been re-consumed as an answer.ProviderRuntimeIngestion.tscarried twocase "turn.aborted"entries: upstream's→ "interrupted"(fix(server): unblock OpenCode approvals and stop pingdotgg/t3code#9653) and a fork-only→ "ready"(shadowed dead code). The fork test pinning"ready"had been failing on main since the merge; the shadow made the live behavior unreadable.Fix
error/stopped/interruptedsettles the ask asfailedwith a status-specific reason (or the write'slastErrorwhen it carries one). The reactor's existing bounded re-drive then fires; when the budget is spent the run fails with the provider's reason. Only aready/idlewrite — the turn actually completed — may settle with text. A dead turn's streamed text is preamble, never an answer.findCompletedAnswerskips messages belonging to the thread's latest turn when that turn endedinterrupted/error(the projection settles dead turns exactly that way). A completed answer from an earlier turn still counts.case "turn.aborted": return "ready"; the live"interrupted"case now documents why it is the right status. The stale fork test is updated to pin"interrupted"(matching upstream #9653's OpenCode tests, which already asserted it).Why this is safe
tracker.forgetand stops the run — the step never settles at all. The dead-state rule only engages when the host is still waiting on the turn.thread.turn.resumeexplicitly accepts a thread whose last turn endedinterrupted/error("a partial assistant message may trail the user message, but the reply never completed"). The session-level transient-retry ladder and the host's step-level re-drive coordinate through the decider's one-active-turn invariant.interruptedis not "dead" in the wait rule. An interrupted session is still alive: a strayinterruptedwrite for a turn the watch never saw running does not end the wait (dead statuseserror/stoppedstill end it unconditionally).Tests
t3team-workflowTurnResolution.test.ts:stoppednow settlesfailed(with/without partial text); newinterruptedcases (pre-abort text refused,lastErrorhonored, stray-write guard).t3team-workflowTurnAnswerLookup.test.ts: dead-latest-turn preamble never counts; a completed reply still counts when only a later turn died.t3team-workflowEngineTurnAnswer.integration.test.ts(host-level regression, both dead shapes): a turn that dies mid-stream — or is aborted — rejects theaskAgentstep and re-drives it; the run completes only with the re-driven turn's answer, and the rejection is journaled as a re-drive attempt on the run row. Pre-existing suites (preamble-not-answer, no-text failure, restart re-drive, budget exhaustion) all still pass.ProviderRuntimeIngestion.test.ts: the abort test now pins"interrupted"(was failing on main).Verification
t3team-workflowTurnResolution.test.ts,t3team-workflowTurnAnswerLookup.test.ts,t3team-workflowEngineReactorTasks.test.ts,t3team-workflowEngineTurnAnswer.integration.test.ts,ProviderRuntimeIngestion.test.ts, plus decider/transient-retry/rehydrate/registry/reactor suites — all green.tsgo --noEmitinapps/server: no errors in the changed files (the remaining diagnostics are pre-existing on main, in untouched files).vp linton all changed files: clean.node t3team-additive-guard.mjs: passes.Related
pj/nexi-distributionFix startup V8 OOM: bound the decider's command read model at boot #297 — turn-layer half: a mid-stream provider death emits no terminal event, so the turn only settles when the 10-minute watchdog interrupts it. This PR makes the host robust to that death however it arrives; the turn layer should still emit a terminal event promptly.