Skip to content

[Bug]: V2 Claude run stays "running" forever after the CLI reports result/success (thread stuck on "Thinking") #15197

Description

@Info-Cado

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server (orchestration V2, Claude adapter)

Steps to reproduce

Intermittent. Happened once in about 120 Claude runs in my local history.

  1. Start a new thread with the Claude provider (claude-opus-5-5[1m], bypassPermissions) in a worktree.
  2. Send a normal coding prompt. The agent runs about 9 tool turns, including a foreground Bash command (tests, build, lint) and a git commit, then writes its final answer.

Expected behavior

When the Claude CLI emits result / subtype: success / terminal_reason: completed, the run moves to completed and the thread goes idle.

Actual behavior

The final assistant message renders, but the thread keeps showing "Thinking" and the "Working for Xm Ys" timer keeps counting indefinitely.

The provider log shows the turn ended cleanly at 2026-10-03T13:40:46.593Z. The last two events are:

{"type":"result","subtype":"success","stop_reason":"end_turn","terminal_reason":"completed","is_error":false,"num_turns":9,"duration_ms":49562,"queued_turn_count":0,"result_index":0, ...}
{"type":"command_lifecycle","command_uuid":"ed0f6284-...","state":"completed", ...}

Server state in statev2.sqlite, checked more than 20 minutes later:

  • orchestration_v2_projection_runs: run:thread:<id>:ordinal:1 has status = running and completed_at = NULL.
  • orchestration_v2_projection_run_attempts: attempt 1 has status = running.
  • orchestration_v2_projection_provider_turns: the provider turn has status = running.
  • orchestration_v2_projection_provider_sessions: the session has status = ready.
  • orchestration_v2_effect_outbox: no pending or failed rows. No checkpoint.capture was ever enqueued for this run, so this is not the waiting / stalled-capture case from [BUG] V2 waiting run keeps showing Agent is working and offers Stop for finished work when checkpoint capture stalls #15124.

Other details:

  • The run had one task_started / task_notification pair (local_bash, is_backgrounded: false, status: completed) at 13:40:38, about 8 s before the result.
  • One command_execution turn item ended as failed at 13:40:18.
  • Neither looks unusual: other runs in the same session with the same patterns completed normally.
  • Server trace spans in the 13:40:44–13:40:50 window all exited with Success. I could not find an error.

The provider session stays ready while the run, attempt and provider turn stay running. That suggests the adapter received the result but the run-terminal transition was never committed.

Impact

Major degradation or frequent failure. The thread looks busy forever, and it is unclear whether the agent is still working.

Version or commit

0.0.46-nightly.20261003.2623 (fed41fa88bb2), the latest nightly at the time of writing.

Environment

Desktop (AppImage), Linux (Arch-based, Hyprland), Claude provider via Claude Agent SDK, orchestration V2.

Logs or stack traces

The provider event log and the server trace for the thread are available on request. I'm not attaching them because they contain prompt text and local paths.

Screenshots, recordings, or supporting files

No response

Workaround

Press Stop, or restart the app.

Related issues and PRs

Activity

  1. Info-Cado commented on Oct 3, 2026

    @Info-Cado
    Author

    Update: Stop does not recover the thread.

    Stop was pressed 7 times at 14:05:15–14:05:16Z and again at 14:09:21Z. Each press:

    • dispatched run.interrupt, which returned Success;
    • enqueued a provider-turn.interrupt effect in orchestration_v2_effect_outbox, which ended succeeded on attempt 1;
    • ran ClaudeAdapterV2.interruptTurn, which returned Success;
    • also dispatched thread.background-work.settle.

    After all of that, the run is still running with completed_at = NULL. The Claude CLI process for this thread had already exited, so I think the interrupt has no live turn to cancel. No terminal event ever comes back, and the run never leaves running.

    The user also has a follow-up message queued on this thread. It now sits behind the stuck run and is never sent.

    I haven't yet tested whether restarting the app clears the run.

  2. juliusmarminge commented on Oct 3, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks @Info-Cado for checking the database state and separating this from the nearby issues. That made this much easier to pin down. From the source, this looks like a real server-side stuck run on the V2 Claude path, not a client refresh glitch and not the checkpoint stall in #15124.

    What the rows tell us

    A successful Claude result only leaves running once ClaudeAdapterV2 finalizes the turn. That emits provider_turn.updated and turn.terminal, and writeFinalRunEvents then saves the run as waiting and enqueues checkpoint.capture (RunExecutionService.ts, around the persistedStatus assignment). Your run, attempt, and provider turn are all still running, and no checkpoint.capture was ever enqueued, so that transition never happened. The provider session's ready status is its default, so it doesn't show the turn finished.

    Where the result could have been dropped

    Seeing result / subtype: success in the protocol log only proves the adapter read the frame. It can still decline to settle the turn. While the prompt echo is unconfirmed, handleSdkMessage treats a result as belonging to another turn when its user_message_uuids don't include this prompt, or when it has no uuid, promptEchoMode is no longer unknown, and origin.kind isn't human (isClaudeResultForOtherTurn). That path logs orchestration-v2.claude-result-for-other-turn and returns without calling finalizeActiveTurn. The assistant text is already rendered by then, which fits a finished answer under a timer that never stops. The command_lifecycle completed frame afterwards doesn't settle anything, and the foreground local_bash task isn't background work, so it doesn't explain this.

    Stop can hit the same gap. provider-turn.interrupt calls interruptTurn with requestRuntimeRestart: true. If the in-memory active turn is already gone, it returns success without emitting turn.terminal, so the outbox row can succeed while the run stays running and any queued follow-up stays blocked behind it.

    Related

    What would help

    A few fields from the logs would show which branch dropped the result. Please redact prompt text and home paths:

    • From the result frame: origin and user_message_uuids / user_message_uuid.
    • Whether any earlier command_lifecycle for this prompt uuid arrived before the result.
    • Whether the server trace has orchestration-v2.claude-result-for-other-turn or orchestration-v2.claude-task-notification-result-dropped around 2026-10-03T13:40:46Z.

    If you restart the app, it would also help to know whether that run row is still running afterwards.

    A maintainer will decide on the fix direction.

  3. added
    bugSomething is broken or behaving incorrectly.
    needs more infoInitial triage showed no bug. Awaiting more info
    via-triageFiled through npx t3 triage
    on Oct 3, 2026
  4. Info-Cado commented on Oct 3, 2026

    @Info-Cado
    Author

    Thanks @juliusmarminge. Here is what I could get. It comes from the provider event log and statev2.sqlite. Prompt text and home paths are redacted.

    1. The result frame

    {
      "type": "result",
      "subtype": "success",
      "uuid": "cfc2ef2d-…",
      "user_message_uuids": ["ed0f6284-…"],
      "user_message_uuid": "ed0f6284-…",
      "num_turns": 9,
      "queued_turn_count": 0,
      "result_index": 0,
      "stop_reason": "end_turn",
      "terminal_reason": "completed",
      "is_error": false
    }
    • The frame has no origin field. I checked the whole raw payload.
    • ed0f6284-… is the uuid of this turn's prompt.offer (13:39:55.230Z). So the result names this turn's own prompt.

    2. command_lifecycle before the result

    Yes. Both arrived about 50 s before the result:

    Time (Z) Frame
    13:39:55.230 outgoing prompt.offer, uuid ed0f6284-…
    13:39:55.991 command_lifecycle queued, command_uuid: ed0f6284-…
    13:39:57.028 command_lifecycle started, command_uuid: ed0f6284-…
    13:39:57.054 system/init
    13:40:00.844 → system/thinking_tokens frames carrying user_message_uuid: ed0f6284-…
    13:40:38.376 / .848 task_started / task_notification (local_bash, is_backgrounded: false, completed)
    13:40:46.593 result (above)
    13:40:46.617 command_lifecycle completed, command_uuid: ed0f6284-…

    No user replay frame echoed the prompt. The only user frames are tool results.

    3. Server trace

    I can't check this any more. The server trace has rotated, and the oldest span left is from about 18:25Z, hours after 13:40:46Z. orchestration_v2_events is also empty, so the DB has no event-level history to fall back on. Sorry.

    New evidence: the adapter had already cleared its active turn

    I missed this in the first report. orchestration_v2_effect_outbox also has a provider-turn.steer row. The queued follow-up was sent at 13:51:20Z, about 10 minutes after the result. It was routed as a steer into run 1, and all 5 attempts failed:

    OrchestrationEffectExecutionError
      [cause]: ProviderTurnControlError
        [cause]: ProviderAdapterSteerRunError: Failed to steer active run provider-turn:…ordinal%253A1%3Aattempt%3A1 on claudeAgent provider thread …
          [cause]: ProviderAdapterProtocolError: claudeAgent provider protocol error: Claude provider turn provider-turn:…ordinal%253A1%3Aattempt%3A1 is not the active turn.
    

    At fed41fa88bb2, steerTurn returns "has no live query" when queryContext is null. So the CLI query was still live at 13:51. It returns "is not the active turn" when activeTurn does not hold this provider turn. As far as I can tell, activeTurn is written in only two places:

    • Ref.set(activeTurn, context) in startTurn. No other provider-turn.start ran on this thread between 13:39:55Z and 14:10:14Z.
    • The Ref.update(activeTurn, …→ null) at the end of finalizeActiveTurn. It only runs after the provider_turn.updated, provider_thread.updated and turn.terminal emits have returned.

    If I'm reading that right, the adapter did finalize this turn. The result was not dropped by isClaudeResultForOtherTurn, and it was not the task-notification drop either. Both of those return early and leave activeTurn set, so the steer would have found the turn. Yet the provider turn, attempt and run all stayed running, and no checkpoint.capture was ever enqueued. That points to a loss after emitProviderEvent, somewhere between the adapter and the projection or writeFinalRunEvents, rather than in the adapter's result handling.

    This would also explain the lost follow-up. The run still looked running, so the message went to the adapter as a steer, and a turn that is already finalized can't accept one.

    After restarting the app

    The app restarted at about 14:10:07Z to install an update (now 0.0.46-nightly.20261003.2632). After that:

    • Run ordinal 1 became cancelled, with completed_at = 14:10:07.924Z. Its provider turn is cancelled too.
    • A provider-runtime.continue effect (sourceRunId = run 1) started run ordinal 2 on a resumed Claude session, with the prompt "Continue where you left off." It completed normally at 14:10:32Z and enqueued checkpoint.capture.
    • The queued follow-up was never delivered. Its text doesn't appear in any prompt.offer in the provider log. It is still stored as a user message attached to run 1.
    • The next message (run ordinal 3) completed normally.

    So a restart clears the stuck run, but the queued follow-up is silently lost.

    Correction to my earlier update

    I said the Claude CLI process had already exited before Stop. That's wrong. The outbox has 16 provider-turn.interrupt rows between 14:05:08Z and 14:05:16Z, plus one at 14:09:21Z. The first one was created at 14:05:08.182Z, and the provider log shows the outgoing query.close at 14:05:08.196Z. So the first Stop is what closed the CLI. The later presses found nothing to cancel.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.needs more infoInitial triage showed no bug. Awaiting more infovia-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions