Repository navigation
[Bug]: V2 Claude run stuck running forever after mid-run provider error — settle/interrupt refused, startup reconciliation never fires #16449
Description
Activity
Note
Grok responding on behalf of Julius.
Triage
Thanks @amihos for the detailed projection state, it made this much easier to trace. From the code, this looks like a real stuck run on
0.0.46-nightly.20261005.2676, but not a startup-reconciliation regression of #4561 and not the same failure as #15197 or #15173.What I found
- The "Provider recovered" item didn't end the run. Position 261 (
Claude API unknown, titledProvider recovered) is a Claudeapi_retryframe whose error name wasunknown. The adapter records it as a retry, then marks the same item recovered when the next assistant message arrives. It isn't a terminal provider failure, so the run isn't expected to move tofailedthere, and it kept going until the completedcommand_executionat2026-10-06T03:00:32.686Z. - Settle and archive. While a run is still
running,thread.settletreats it as active work and rejects with the message you saw. Archive only setsarchivedand doesn't stop the run. - Stop is where it gets stuck.
Failed to dispatch orchestration command run.interruptmeans the decider rejected Stop, not that the Claude process ignored an interrupt. On.2676, when the provider turn is stillrunningbut the in-memory provider session is already gone,run.interruptfails (with eitherProvider session … is not active.orProvider thread … has no active provider session.; MCP only shows the outer sentence, see [Bug]: MCP thread tools replace the orchestrator's rejection reason with "The operation could not be completed." #15586) instead of ending the run. The follow-upthread.background-work.settlealso returns early while any run is stillrunning. - Startup reconciliation didn't run against it. The fix(server): reconcile orphaned provider sessions #7719 reconciliation cancels
runningruns only when the server process starts or shuts down. The run still has its originalstartedAtandcompletedAt: null, which fits a server that was never restarted overnight. - fix(orchestration-v2): let Stop recover stalled runs #15442 on
mainchanges this path. It merged after.2676. When the session is already gone, Stop settles the matching stalled attempt locally, keeps partial output, and captures the checkpoint, so it should give a way out once it's in a nightly.
Workaround for now: restarting the app or server should let startup reconciliation mark the run
cancelled(other active threads on that server get interrupted too), after which settle should work.Likely fix area
The open question is why the session disappeared without a terminal status being committed. Two known ways to end up here:
- The run's event consumer died on a failed write and the idle session was released about 30 minutes later, the path tracked in open fix(server): V2 runs no longer look busy forever after their events stop saving #14856.
- The provider's terminal status never committed, like the disk-full case fix(orchestration-v2): let Stop recover stalled runs #15442 describes.
If you're able to share a few redacted lines from
server.trace.ndjson, these would show which one it was:- the
causeon therun.interruptOrchestratorDispatchError - the provider session row's status and
updated_at - any
database is locked, failed event write,orchestration-v2.claude-query-stream-failed, ororchestration-v2.driver-session.idle-releaseafter the last timeline item
A maintainer will decide on the fix direction.
- The "Provider recovered" item didn't end the run. Position 261 (
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Oct 6, 2026 Thanks for tracing this — answers to the three asks:
1. The
causeon therun.interruptOrchestratorDispatchErrorReproduced fresh just now via
t3_thread_interruptand captured fromserver.trace.ndjson, spanorchestrationV2.dispatch.once:[cause]: Error: Provider session provider-session:provider-instance:claudeAgent_stock:thread:6424679f-…:f17d643f-… is not active.So it's your first variant (
…is not active.), hitting the path where the provider turn is stillrunningbut the in-memory session is gone. (Side note confirming #15586: the MCP tool result only carried the outerFailed to dispatch orchestration command run.interruptsentence — the real cause was only in the trace.)2. Provider session row status /
updated_atFrom
statev2.sqlite,orchestration_v2_projection_provider_sessionsfor this thread (single row):status = stopped,updated_at = 2026-10-06T08:02:06.486Z, modelclaude-opus-5-5
Against that:
- run
ordinal:10:status = running,requested_at 2026-10-06T02:50:04.227Z,completed_at = NULL - provider turn
ordinal:10:status = running,started_at 2026-10-06T02:50:04.720Z,completed_at = NULL - last timeline item (pos 270,
command_execution,completed):2026-10-06T03:00:32.686Z
So the session row went
stopped~5h after the last timeline activity — much later than the ~30 min idle-release window, which seems worth noting. I did not touch this thread between 03:00 and discovery (was away from the PC); the only writes since were my triage just now (queue-cancel of orphaned run 11, archive).3.
database is locked/ failed event write /claude-query-stream-failed/driver-session.idle-releaseafter the last timeline itemCan't provide these: local
server.trace.ndjson*retention only covers roughly the last hour (11 rotated files, ~11 MB each, earliest coverage ~11:46 local today). Nothing from the 03:00–08:02 window survives locally. In the retained window there are zero hits fordatabase is locked,query-stream/stream-failed,idle-release/idleRelease,no active provider session, orhas active or blocked work.No-restart confirmed
Server process (pid 3938751) started
Mon Oct 5 15:50:09local (server-runtime.json startedAt 2026-10-05T14:50:14.739Z), uptime 21h at check — so no restart happened overnight and #7719 reconciliation indeed never ran against this run, matching your reading.I have not restarted yet since that would interrupt other active threads on this server. Happy to do it if you want the reconciliation confirmation, or to wait for the #15442 nightly and test Stop then — just say which.
Before submitting
Area
apps/server (orchestration V2, Claude adapter)
What happened
A V2 Claude thread is permanently stuck
runningand cannot be settled, stopped, or interrupted — after a mid-run provider error, not a restart. Looks like a regression / uncovered path of #4561 (closed by #7719, which only reconciles at startup).Thread ID:
6424679f-b22f-4795-ab90-678559fae04cT3 Code version:
0.0.46-nightly.20261005.2676Provider:
claudeAgent_stock/claude-opus-5-5,runtimeMode: full-accessState via thread API:
thread.status:running,settled: falseactiveRunId:run:thread:6424679f-...:ordinal:10, stuckrunningrequestedAt:2026-10-06T02:50:04.227Z,startedAt:2026-10-06T02:50:04.272Z,completedAt: nullcommand_execution,completed,updatedAt 2026-10-06T03:00:32.686Z(~8h silence, stillrunning)error: "Claude API unknown", title "Provider recovered". The run continued for ~10 min after that, then went silent without ever transitioning tofailed/completed.pendingRequestCount: 0, no pending user inputs,transfers: []— not the AskUserQuestion-leak variant ([Bug]: ClaudeAdapter leaks pending AskUserQuestion requests on session teardown — thread becomes unanswerable and unsettleable #5119).ordinal:11, "Background command Wait for CI...") existed; cancelling it via API succeeded and the queue is now empty, but settle is still refused.failed(Claude API rate limit reachedafter a usage-limit pause notice). Run 12:cancelled.On restarts: I was away from the PC overnight and did not manually restart anything. I can't rule out an unattended nightly auto-update between 03:00 and discovery.
Expected behavior
A provider error with no live session behind the run should move the run to
failed(or startup/interrupt reconciliation should end it), so the thread becomes settleable.Actual behavior
Thread 6424679f-b22f-4795-ab90-678559fae04c has active or blocked work and cannot be settled.t3_thread_interrupt(with and without explicitrunIdfor ordinal 10) →Failed to dispatch orchestration command run.interruptt3_thread_organizesettle via MCP →{"_tag":"OrchestratorMcpFailure","code":"orchestration_error","message":"The operation could not be completed."}(the generic-masking bug in [Bug]: MCP thread tools replace the orchestrator's rejection reason with "The operation could not be completed." #15586 — the real reason only shows in the UI/server trace)t3_thread_organizearchive → succeeded (archived: true), butstatusis stillrunning/settled: falseunderneath. Archive hides it; doesn't fix it.Related