Skip to content

[Bug]: V2 Claude run stuck running forever after mid-run provider error — settle/interrupt refused, startup reconciliation never fires #16449

Description

@amihos

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server (orchestration V2, Claude adapter)

What happened

A V2 Claude thread is permanently stuck running and cannot be settled, stopped, or interrupted — after a mid-run provider error, not a restart. Looks like a regression / uncovered path of #4561 (closed by #7719, which only reconciles at startup).

Thread ID: 6424679f-b22f-4795-ab90-678559fae04c
T3 Code version: 0.0.46-nightly.20261005.2676
Provider: claudeAgent_stock / claude-opus-5-5, runtimeMode: full-access

State via thread API:

  • thread.status: running, settled: false
  • activeRunId: run:thread:6424679f-...:ordinal:10, stuck running
    • requestedAt: 2026-10-06T02:50:04.227Z, startedAt: 2026-10-06T02:50:04.272Z, completedAt: null
    • last timeline item: position 270, command_execution, completed, updatedAt 2026-10-06T03:00:32.686Z (~8h silence, still running)
  • Position 261 in the same run: error: "Claude API unknown", title "Provider recovered". The run continued for ~10 min after that, then went silent without ever transitioning to failed/completed.
  • pendingRequestCount: 0, no pending user inputs, transfers: [] — not the AskUserQuestion-leak variant ([Bug]: ClaudeAdapter leaks pending AskUserQuestion requests on session teardown — thread becomes unanswerable and unsettleable #5119).
  • One orphaned queued run (ordinal:11, "Background command Wait for CI...") existed; cancelling it via API succeeded and the queue is now empty, but settle is still refused.
  • Previous run 9: failed (Claude API rate limit reached after a usage-limit pause notice). Run 12: cancelled.

On restarts: I was away from the PC overnight and did not manually restart anything. I can't rule out an unattended nightly auto-update between 03:00 and discovery.

Expected behavior

A provider error with no live session behind the run should move the run to failed (or startup/interrupt reconciliation should end it), so the thread becomes settleable.

Actual behavior

  • Sidebar/UI Settle → Thread 6424679f-b22f-4795-ab90-678559fae04c has active or blocked work and cannot be settled.
  • t3_thread_interrupt (with and without explicit runId for ordinal 10) → Failed to dispatch orchestration command run.interrupt
  • t3_thread_organize settle via MCP → {"_tag":"OrchestratorMcpFailure","code":"orchestration_error","message":"The operation could not be completed."} (the generic-masking bug in [Bug]: MCP thread tools replace the orchestrator's rejection reason with "The operation could not be completed." #15586 — the real reason only shows in the UI/server trace)
  • t3_thread_organize archive → succeeded (archived: true), but status is still running / settled: false underneath. Archive hides it; doesn't fix it.

Related

Activity

  1. juliusmarminge commented on Oct 6, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks @amihos for the detailed projection state, it made this much easier to trace. From the code, this looks like a real stuck run on 0.0.46-nightly.20261005.2676, but not a startup-reconciliation regression of #4561 and not the same failure as #15197 or #15173.

    What I found

    • The "Provider recovered" item didn't end the run. Position 261 (Claude API unknown, titled Provider recovered) is a Claude api_retry frame whose error name was unknown. The adapter records it as a retry, then marks the same item recovered when the next assistant message arrives. It isn't a terminal provider failure, so the run isn't expected to move to failed there, and it kept going until the completed command_execution at 2026-10-06T03:00:32.686Z.
    • Settle and archive. While a run is still running, thread.settle treats it as active work and rejects with the message you saw. Archive only sets archived and doesn't stop the run.
    • Stop is where it gets stuck. Failed to dispatch orchestration command run.interrupt means the decider rejected Stop, not that the Claude process ignored an interrupt. On .2676, when the provider turn is still running but the in-memory provider session is already gone, run.interrupt fails (with either Provider session … is not active. or Provider thread … has no active provider session.; MCP only shows the outer sentence, see [Bug]: MCP thread tools replace the orchestrator's rejection reason with "The operation could not be completed." #15586) instead of ending the run. The follow-up thread.background-work.settle also returns early while any run is still running.
    • Startup reconciliation didn't run against it. The fix(server): reconcile orphaned provider sessions #7719 reconciliation cancels running runs only when the server process starts or shuts down. The run still has its original startedAt and completedAt: null, which fits a server that was never restarted overnight.
    • fix(orchestration-v2): let Stop recover stalled runs #15442 on main changes this path. It merged after .2676. When the session is already gone, Stop settles the matching stalled attempt locally, keeps partial output, and captures the checkpoint, so it should give a way out once it's in a nightly.

    Workaround for now: restarting the app or server should let startup reconciliation mark the run cancelled (other active threads on that server get interrupted too), after which settle should work.

    Likely fix area

    The open question is why the session disappeared without a terminal status being committed. Two known ways to end up here:

    If you're able to share a few redacted lines from server.trace.ndjson, these would show which one it was:

    • the cause on the run.interrupt OrchestratorDispatchError
    • the provider session row's status and updated_at
    • any database is locked, failed event write, orchestration-v2.claude-query-stream-failed, or orchestration-v2.driver-session.idle-release after the last timeline item

    A maintainer will decide on the fix direction.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 6, 2026
  3. amihos commented on Oct 6, 2026

    @amihos
    Author

    Thanks for tracing this — answers to the three asks:

    1. The cause on the run.interrupt OrchestratorDispatchError

    Reproduced fresh just now via t3_thread_interrupt and captured from server.trace.ndjson, span orchestrationV2.dispatch.once:

    [cause]: Error: Provider session provider-session:provider-instance:claudeAgent_stock:thread:6424679f-…:f17d643f-… is not active.

    So it's your first variant (…is not active.), hitting the path where the provider turn is still running but the in-memory session is gone. (Side note confirming #15586: the MCP tool result only carried the outer Failed to dispatch orchestration command run.interrupt sentence — the real cause was only in the trace.)

    2. Provider session row status / updated_at

    From statev2.sqlite, orchestration_v2_projection_provider_sessions for this thread (single row):

    • status = stopped, updated_at = 2026-10-06T08:02:06.486Z, model claude-opus-5-5

    Against that:

    • run ordinal:10: status = running, requested_at 2026-10-06T02:50:04.227Z, completed_at = NULL
    • provider turn ordinal:10: status = running, started_at 2026-10-06T02:50:04.720Z, completed_at = NULL
    • last timeline item (pos 270, command_execution, completed): 2026-10-06T03:00:32.686Z

    So the session row went stopped ~5h after the last timeline activity — much later than the ~30 min idle-release window, which seems worth noting. I did not touch this thread between 03:00 and discovery (was away from the PC); the only writes since were my triage just now (queue-cancel of orphaned run 11, archive).

    3. database is locked / failed event write / claude-query-stream-failed / driver-session.idle-release after the last timeline item

    Can't provide these: local server.trace.ndjson* retention only covers roughly the last hour (11 rotated files, ~11 MB each, earliest coverage ~11:46 local today). Nothing from the 03:00–08:02 window survives locally. In the retained window there are zero hits for database is locked, query-stream/stream-failed, idle-release/idleRelease, no active provider session, or has active or blocked work.

    No-restart confirmed

    Server process (pid 3938751) started Mon Oct 5 15:50:09 local (server-runtime.json startedAt 2026-10-05T14:50:14.739Z), uptime 21h at check — so no restart happened overnight and #7719 reconciliation indeed never ran against this run, matching your reading.

    I have not restarted yet since that would interrupt other active threads on this server. Happy to do it if you want the reconciliation confirmation, or to wait for the #15442 nightly and test Stop then — just say which.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions