Repository navigation
[Bug]: V2 run stays starting forever after provider-turn.start exhausts its retries — no Stop button, settle refused, only Delete clears it (Pi provider, desktop) #13392
Description
Activity
Confirmed on the current V2 head (
t3code/codex-turn-mapping@692ad0a8fd, #2829). This is a real bug, and the two gaps in the report are both still present. The Pi process exiting is local; T3 never terminalizes the run, and desktop cannot interrupt it.Server: a terminal
provider-turn.startfailure does not settle the runEffectOutbox.claimNextincrementsattempt_countbefore execution, andEffectWorkerpasseswillRetry: effect.attemptCount < maxAttempts(default 5). The last attempt therefore runs withwillRetry === false.ProviderTurnStartService.startonly settles whenproviderSessions.openfails andwillRetryis false (settleRunBeforeStart, signalprovider-session-open-failure). That path is tested. Errors after a successful open are not.resumeThreadfailure falls through toensureThread(fork / handoff prep can fail the same way); the error escapesstart, is wrapped asProviderTurnStartError, and no run event is written — even on the final attempt.EffectWorkerthen takesoutbox.fail, which only setsorchestration_v2_effect_outbox.status = 'failed'. The projection run staysstartingand its attempt stayspending.thread.settlerejects that non-terminal run ("has active or blocked work").run.interruptwould already finish it: astartingrun with no provider turn is settled as "Run interrupted before provider start" and the start effect is cancelled. Nothing emits that command from desktop.ProviderRuntimeRecoveryService.reconcile("startup")cancels non-terminal runs, which is why a restart clears the thread and a live process does not.One correction to the timing note: with 5 attempts the backoff is 100 / 200 / 400 / 800 ms. The
30svalue is only the cap in the formula. The ~4 minutes in the trace (06:32:35Z → 06:36:46Z) is five slow Pi launch failures, not the retry delay.Desktop: Stop is hidden for this state
derivePhasemapspreparing/starting/queuedto"connecting", and onlyrunning/waitingto"running". The composer renders Stop only whenisRunning(phase === "running"), and thethread.stopshortcut uses the samecanInterruptRunningThreadgate. While phase stays"connecting", local dispatch is never acknowledged, so the primary control stays a disabled spinner instead of Stop. The sidebar still treatsstartingas working.Mobile is not the same UI bug.
threadRuntimeHasInterruptibleRunalready treatspreparingandstartingas interruptible, so a mobile Stop in this state should reachrun.interrupt. The run still never fails on its own if nobody stops it.Not a duplicate
- [Bug]: Mobile Stop does nothing during 'Setting up worktree…' (preparing phase) #12187 is mobile Stop during worktree
preparing(no session yet). Same missing-Stop family, different phase and client. - [Bug]: Turn interrupt on a stopped session settles nothing, leaving a dangling active turn and a permanently failing stop #11796 is an interrupt against an already stopped session.
- [Bug]: Remote stop button gives no feedback and leaves thread stuck in Thinking #8618 is remote Stop feedback while Thinking.
Fix
Both halves, on the V2 branch:
- When
willRetry === falseand the run is stillstarting, catch post-open failures insideProviderTurnStartService.startand settle with the samefailed+provider_erroritem as a failed session open, then return success so the outbox is not leftfailedon an already-settled run. KeepwillRetry === trueas a retry that leaves the runstarting. If the settle write itself fails, leave the effect retryable (same as the existing open-failure persistence test). - Offer desktop Stop (button and
thread.stop) for an interruptiblepreparingorstartingrun — the same set asthreadRuntimeHasInterruptibleRun. Do not treatqueuedas interruptible;dispatchRunInterruptrejects it. The connecting-phase send spinner has to yield to that Stop button.
Labels:
bug,accepted(keepvia-triage).- [Bug]: Mobile Stop does nothing during 'Setting up worktree…' (preparing phase) #12187 is mobile Stop during worktree
- addedacceptedfeature request acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.
on Sep 24, 2026 Closing permanently: V2-only bug (provider-turn.start / orchestration-v2 not on main). Not actionable on current main; do not reopen.
What happened
On a Pi-provider thread that had already completed six turns, I sent a seventh message. The Pi child process died at turn start, and the thread then sat in the sidebar as working for over two hours. The composer never offered a Stop button, so I could not interrupt it. "Settle thread" failed with:
Failed to settle thread. Thread 77e0f82b-… has active or blocked work and cannot be settled.The only thing that cleared it was Delete thread.The surfaced error (thread detail) was:
ProviderAdapterEventStreamError: Failed while streaming pi provider session provider-session:provider-instance:pi:thread:77e0f82b-…:f2aa6460-… events.Diagnosis
Two gaps combine, and neither depends on why Pi exited (that part is a Pi-side launch failure on my machine; the report is about how T3 handled it).
1. Server — a terminally failed
provider-turn.starteffect never settles its run.EffectWorker.runOnceretries the effect (maxAttempts= 5, backoff up to 30 s). On the last failure it callsoutbox.fail(...), which only flipsorchestration_v2_effect_outbox.statustofailed. No domain event follows, soorchestration_v2_projection_runskeeps the run atstatus: "starting"withstartedAt: nulland its attempt atpending.ProviderTurnStartServicedoes settle the run whenproviderSessions.openfails (settleRunBeforeStart→failed, gated onwillRetry), but an error that escapes after open — heresession.resumeThreadfailed, the fallbacksession.ensureThreadthen failed withProviderAdapterEventStreamError←PiRpcError: Pi RPC read failed: pi process exited with code 1(raised fromPiAdapterV2's transport-death handler) — propagates asProviderTurnStartErrorwith no run settlement at all. The provider session row does record it (status: "error",lastErrorbelow), but the run does not. OnlyProviderRuntimeRecoveryService.reconcile("startup")(an app restart) terminalizes non-terminal runs.2. Client — the Stop affordance is hidden for exactly this state, even though the server would honour it.
derivePhase(apps/web/src/session-logic.ts) mapspreparing/starting/queued→"connecting", andcanInterruptRunningThread(apps/web/src/components/ChatView.tsx) requiresphase === "running". So a run stuck instartingshows a disabled composer and no Stop. MeanwhiledispatchRunInterruptinOrchestrator.tsexplicitly handles astartingrun with no provider turn ("Run interrupted before provider start") — the server trace shows onlythread.settlecommands from my attempts, never arun.interrupt, because the UI never offered one.thread.settlethen refuses because a non-terminal run exists.Net effect: an unrecoverable startup failure leaves a thread permanently "working" with no user-reachable exit except delete or an app restart.
Verified against the current #2829 head (692ad0a):
EffectWorker.ts,EffectOutbox.ts, therun.interrupthandler,derivePhaseandcanInterruptRunningThreadare unchanged from 060756d.Steps to reproduce
resumeThread).pishim first on PATH that doesexit 1, or drop a global extension that throws on import.Expected: after the retry budget the run fails with a visible provider error; while it is retrying, Stop is available and cancels it; the thread can be settled afterwards.
Actual: ~4 min of retries (5 attempts), then the outbox effect is
failed; the run staysstarting/ attemptpendingindefinitely; the composer shows no Stop;thread.settle→ "has active or blocked work"; only Delete (or an app restart) clears it.Version
0.0.42 — V2 branch (#2829) at 060756d, personal packaged build; code paths verified identical at 692ad0a
Environment
macOS 26.6.2 (arm64), Electron 44.4.2; Pi 0.87.1 on Node 26.8.2 (Homebrew); local desktop client
Evidence
Related issues
#12187 (mobile Stop is a no-op during the
preparing/worktree phase — same hidden-Stop class, different phase and client; this one is desktop + a server-side run that is never settled). #11796 and #8618 concern interrupts on running/stopped sessions, not a run that never started.Fix applied or workaround
Nothing was changed on the machine. Unblocked by deleting the thread; restarting the app would also have cleared it (startup reconcile cancels non-terminal runs). Suggested fix, both halves: (a) when
EffectWorkertakes the terminaloutbox.failbranch forprovider-turn.start, settle the run asfailedwith a provider_error item (the same shapesettleRunBeforeStartproduces for a failed session open) — or haveProviderTurnStartService.startcatch escaping errors whenwillRetry === falseand do that itself; (b) let the composer offer Stop while the active run ispreparing/starting(phase"connecting"), sincerun.interruptalready supports that state.Filed by
Pi (claude-opus-5-5) driven by Alex Stark (astarktc); investigation from the live DB + server trace + fork source, filed through the triage form