Skip to content

[Bug]: V2 run stays starting forever after provider-turn.start exhausts its retries — no Stop button, settle refused, only Delete clears it (Pi provider, desktop) #13392

Description

@astarktc

What happened

On a Pi-provider thread that had already completed six turns, I sent a seventh message. The Pi child process died at turn start, and the thread then sat in the sidebar as working for over two hours. The composer never offered a Stop button, so I could not interrupt it. "Settle thread" failed with: Failed to settle thread. Thread 77e0f82b-… has active or blocked work and cannot be settled. The only thing that cleared it was Delete thread.

The surfaced error (thread detail) was:
ProviderAdapterEventStreamError: Failed while streaming pi provider session provider-session:provider-instance:pi:thread:77e0f82b-…:f2aa6460-… events.

Diagnosis

Two gaps combine, and neither depends on why Pi exited (that part is a Pi-side launch failure on my machine; the report is about how T3 handled it).

1. Server — a terminally failed provider-turn.start effect never settles its run. EffectWorker.runOnce retries the effect (maxAttempts = 5, backoff up to 30 s). On the last failure it calls outbox.fail(...), which only flips orchestration_v2_effect_outbox.status to failed. No domain event follows, so orchestration_v2_projection_runs keeps the run at status: "starting" with startedAt: null and its attempt at pending. ProviderTurnStartService does settle the run when providerSessions.open fails (settleRunBeforeStart → failed, gated on willRetry), but an error that escapes after open — here session.resumeThread failed, the fallback session.ensureThread then failed with ProviderAdapterEventStreamError ← PiRpcError: Pi RPC read failed: pi process exited with code 1 (raised from PiAdapterV2's transport-death handler) — propagates as ProviderTurnStartError with no run settlement at all. The provider session row does record it (status: "error", lastError below), but the run does not. Only ProviderRuntimeRecoveryService.reconcile("startup") (an app restart) terminalizes non-terminal runs.

2. Client — the Stop affordance is hidden for exactly this state, even though the server would honour it. derivePhase (apps/web/src/session-logic.ts) maps preparing/starting/queued → "connecting", and canInterruptRunningThread (apps/web/src/components/ChatView.tsx) requires phase === "running". So a run stuck in starting shows a disabled composer and no Stop. Meanwhile dispatchRunInterrupt in Orchestrator.ts explicitly handles a starting run with no provider turn ("Run interrupted before provider start") — the server trace shows only thread.settle commands from my attempts, never a run.interrupt, because the UI never offered one. thread.settle then refuses because a non-terminal run exists.

Net effect: an unrecoverable startup failure leaves a thread permanently "working" with no user-reachable exit except delete or an app restart.

Verified against the current #2829 head (692ad0a): EffectWorker.ts, EffectOutbox.ts, the run.interrupt handler, derivePhase and canInterruptRunningThread are unchanged from 060756d.

Steps to reproduce

  1. Pi provider thread with at least one completed turn (so the next turn goes through resumeThread).
  2. Make the Pi child fail at launch deterministically — e.g. put a pi shim first on PATH that does exit 1, or drop a global extension that throws on import.
  3. Send a message.

Expected: after the retry budget the run fails with a visible provider error; while it is retrying, Stop is available and cancels it; the thread can be settled afterwards.

Actual: ~4 min of retries (5 attempts), then the outbox effect is failed; the run stays starting / attempt pending indefinitely; the composer shows no Stop; thread.settle → "has active or blocked work"; only Delete (or an app restart) clears it.

Version

0.0.42 — V2 branch (#2829) at 060756d, personal packaged build; code paths verified identical at 692ad0a

Environment

macOS 26.6.2 (arm64), Electron 44.4.2; Pi 0.87.1 on Node 26.8.2 (Homebrew); local desktop client

Evidence

# orchestration_v2_projection_runs — the stuck run (ordinal 7); ordinals 1–6 completed normally
run_id=run:thread:77e0f82b-…:ordinal:7  status=starting  requested_at=2026-09-24T06:32:35.405Z  startedAt=null  completedAt=null
  activeAttemptId=…:ordinal:7:attempt:1  (orchestration_v2_projection_run_attempts.status=pending, startedAt=null)
# ordinal 8 (my retry) was queued and cancelled 06:48:26Z; thread updated_at stays 06:48:26Z from then on

# orchestration_v2_effect_outbox — the effect was retried 5x over 4 min and then marked failed; nothing else follows
effect_id=effect:aaac8de6-…:provider-turn.start:run:thread:77e0f82b-…:ordinal:7
  status=failed  attempt_count=5  created_at=2026-09-24T06:32:35.416Z  completed_at=2026-09-24T06:36:46.643Z
  last_error=OrchestrationEffectExecutionError:
    at file:///Applications/T3 Code (Alpha).app/Contents/Resources/app.asar/apps/server/dist/binCli-CnI_s3fV.mjs:213193:32
    at ServerRuntimeStartup.startEffectWorkerWithRelay (Alpha)
    at server.startup …

# orchestration_v2_projection_provider_sessions — the provider session recorded the failure; the run did not
provider_session_id=provider-session:provider-instance:pi:thread:77e0f82b-…:f2aa6460-…
  status=error  createdAt=2026-09-24T06:36:04.366Z  updatedAt=2026-09-24T06:36:46.646Z
  lastError=ProviderAdapterEventStreamError: Failed while streaming pi provider session provider-session:provider-instance:pi:thread:77e0f82b-…:f2aa6460-… events.
    at …/binCli-CnI_s3fV.mjs:190576:28
    at orchestrationV2.providerTurnStart.start (…/binCli-CnI_s3fV.mjs:212954:59)
    at ServerRuntimeStartup.startEffectWorkerWithRelay (Alpha)
    at server.startup.orchestration-v2.effect-worker.start (…/binCli-CnI_s3fV.mjs:223952:99) {
    [cause]: Error: Cause([Fail(PiRpcError: Pi RPC read failed: pi process exited with code 1.)])
  }

# server.trace.ndjson — every command my UI attempts produced was thread.settle; no run.interrupt was ever dispatched
orchestrationV2.dispatch.threadMutation  exit=Failure  cause=OrchestratorDispatchError: Failed to dispatch orchestration command thread.settle (f570f292-…)   startTime=2026-09-24T08:37:44Z
orchestrationV2.dispatch.threadMutation  exit=Failure  cause=OrchestratorDispatchError: Failed to dispatch orchestration command thread.settle (35bf3964-…)   startTime=2026-09-24T08:38:19Z
orchestrationV2.dispatch.threadMutation  exit=Failure  cause=OrchestratorDispatchError: Failed to dispatch orchestration command thread.settle (eac082c4-…)   startTime=2026-09-24T08:40:55Z
orchestrationV2.dispatch.threadMutation  exit=Failure  cause=OrchestratorDispatchError: Failed to dispatch orchestration command thread.settle (104f6ebb-…)   startTime=2026-09-24T08:41:14Z

# source (060756de5a == 692ad0a8fd for these)
apps/server/src/orchestration-v2/EffectWorker.ts ~L645   effect.attemptCount >= maxAttempts → outbox.fail({effectId, workerId, error})   (no run settlement)
apps/server/src/orchestration-v2/EffectOutbox.ts  L589   fail: UPDATE orchestration_v2_effect_outbox SET status='failed' …
apps/server/src/orchestration-v2/Orchestrator.ts  dispatchRunInterrupt: providerTurn===undefined && run.status in (preparing|starting|running) → "Run interrupted before provider start"
apps/web/src/session-logic.ts  derivePhase: starting → "connecting"
apps/web/src/components/ChatView.tsx  canInterruptRunningThread = activeThread !== undefined && phase === "running"

Related issues

#12187 (mobile Stop is a no-op during the preparing/worktree phase — same hidden-Stop class, different phase and client; this one is desktop + a server-side run that is never settled). #11796 and #8618 concern interrupts on running/stopped sessions, not a run that never started.

Fix applied or workaround

Nothing was changed on the machine. Unblocked by deleting the thread; restarting the app would also have cleared it (startup reconcile cancels non-terminal runs). Suggested fix, both halves: (a) when EffectWorker takes the terminal outbox.fail branch for provider-turn.start, settle the run as failed with a provider_error item (the same shape settleRunBeforeStart produces for a failed session open) — or have ProviderTurnStartService.start catch escaping errors when willRetry === false and do that itself; (b) let the composer offer Stop while the active run is preparing/starting (phase "connecting"), since run.interrupt already supports that state.

Filed by

Pi (claude-opus-5-5) driven by Alex Stark (astarktc); investigation from the live DB + server trace + fork source, filed through the triage form

Activity

  1. juliusmarminge commented on Sep 24, 2026

    @juliusmarminge
    Member

    Confirmed on the current V2 head (t3code/codex-turn-mapping @ 692ad0a8fd, #2829). This is a real bug, and the two gaps in the report are both still present. The Pi process exiting is local; T3 never terminalizes the run, and desktop cannot interrupt it.

    Server: a terminal provider-turn.start failure does not settle the run

    EffectOutbox.claimNext increments attempt_count before execution, and EffectWorker passes willRetry: effect.attemptCount < maxAttempts (default 5). The last attempt therefore runs with willRetry === false.

    ProviderTurnStartService.start only settles when providerSessions.open fails and willRetry is false (settleRunBeforeStart, signal provider-session-open-failure). That path is tested. Errors after a successful open are not. resumeThread failure falls through to ensureThread (fork / handoff prep can fail the same way); the error escapes start, is wrapped as ProviderTurnStartError, and no run event is written — even on the final attempt.

    EffectWorker then takes outbox.fail, which only sets orchestration_v2_effect_outbox.status = 'failed'. The projection run stays starting and its attempt stays pending. thread.settle rejects that non-terminal run ("has active or blocked work"). run.interrupt would already finish it: a starting run with no provider turn is settled as "Run interrupted before provider start" and the start effect is cancelled. Nothing emits that command from desktop. ProviderRuntimeRecoveryService.reconcile("startup") cancels non-terminal runs, which is why a restart clears the thread and a live process does not.

    One correction to the timing note: with 5 attempts the backoff is 100 / 200 / 400 / 800 ms. The 30s value is only the cap in the formula. The ~4 minutes in the trace (06:32:35Z → 06:36:46Z) is five slow Pi launch failures, not the retry delay.

    Desktop: Stop is hidden for this state

    derivePhase maps preparing / starting / queued to "connecting", and only running / waiting to "running". The composer renders Stop only when isRunning (phase === "running"), and the thread.stop shortcut uses the same canInterruptRunningThread gate. While phase stays "connecting", local dispatch is never acknowledged, so the primary control stays a disabled spinner instead of Stop. The sidebar still treats starting as working.

    Mobile is not the same UI bug. threadRuntimeHasInterruptibleRun already treats preparing and starting as interruptible, so a mobile Stop in this state should reach run.interrupt. The run still never fails on its own if nobody stops it.

    Not a duplicate

    Fix

    Both halves, on the V2 branch:

    1. When willRetry === false and the run is still starting, catch post-open failures inside ProviderTurnStartService.start and settle with the same failed + provider_error item as a failed session open, then return success so the outbox is not left failed on an already-settled run. Keep willRetry === true as a retry that leaves the run starting. If the settle write itself fails, leave the effect retryable (same as the existing open-failure persistence test).
    2. Offer desktop Stop (button and thread.stop) for an interruptible preparing or starting run — the same set as threadRuntimeHasInterruptibleRun. Do not treat queued as interruptible; dispatchRunInterrupt rejects it. The connecting-phase send spinner has to yield to that Stop button.

    Labels: bug, accepted (keep via-triage).

  2. added
    acceptedfeature request accepted
    bugSomething is broken or behaving incorrectly.
    on Sep 24, 2026
  3. macodev00 commented on Sep 29, 2026

    @macodev00
    Contributor

    Closing permanently: V2-only bug (provider-turn.start / orchestration-v2 not on main). Not actionable on current main; do not reopen.

  4. juliusmarminge commented on Sep 30, 2026

    @juliusmarminge
    Member

    Fixed by #14183 (server) + #14201 (Stop while starting), both merged into t3code/codex-turn-mapping.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions