Skip to content

[Bug]: Stuck OpenCode deletion blocks new chats in ThreadDeletionReactor.drainThrough #11889

Description

@AC40

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to investigate the problem.

Area

apps/server

Summary

Deleting an OpenCode thread left processThreadDeleted stuck in cancelPendingOpenCodePrompt for 242,562,870 ms, about 67.4 hours. New thread creation then blocked in ThreadDeletionReactor.drainThrough, although the remote environment remained connected and HTTP requests succeeded.

Steps to reproduce

Observed sequence:

  1. Run a headless T3 0.0.40 server on Linux with the OpenCode provider, connected to the macOS desktop client.
  2. Delete an OpenCode thread while its prompt admission/cancellation state is unsettled.
  3. Create a new chat and send its first message.
  4. The thread.created event is persisted, but the request waits in ThreadDeletionReactor.drainThrough and never reaches thread.turn-start-requested.

The exact timing that caused the original OpenCode cancellation to remain unsettled has not been reproduced deterministically. The blocked new-thread path was reproduced twice against the affected live process with the following harmless probe:

  • Read an existing accepted thread.created command ID, payload, and receipt sequence from SQLite in read-only mode.
  • Replay it as a thread.create command over authenticated orchestration.dispatchCommand.
  • T3 deduplicates the existing command ID, then still executes the thread-creation drain barrier.
  • Assert that the RPC acknowledges the existing receipt within eight seconds.

This probe creates no additional thread and submits no model prompt. It timed out before recovery and passed in 18 ms after restarting the primary server.

Expected behavior

Provider teardown should finish or fail within a bounded period. A stuck deletion should not prevent unrelated new chats from starting.

Actual behavior

The client remains connected. New chat requests persist the empty thread, then wait indefinitely behind the shared deletion worker. Two original dispatch spans lasted 121,844.8 ms and 47,730.2 ms before client interruption. On service shutdown, the previously open deletion span finally appeared with a duration of 242,562,870 ms.

Impact

Blocks new chat creation on the affected server until recovery.

Version or commit

T3 server and desktop 0.0.40. The unbounded cancellation wait and shared drain are also present in the upstream main files inspected on September 15, 2026.

Environment

Linux x86_64 headless server, Node v22.23.2, macOS T3 desktop 0.0.40, OpenCode provider. Private paths, thread IDs, and chat content omitted.

Logs or stack traces

Sanitized summary of the shutdown trace:

span: processThreadDeleted
durationMs: 242562870.251242
exit: Interrupted during server shutdown

cancelPendingOpenCodePrompt
  <- stopOpenCodeContext
  <- OpenCodeAdapter.stopSession
  <- ProviderService.stopSession
  <- processThreadDeleted

new-chat request:
ws.rpc.orchestration.dispatchCommand
  -> ThreadDeletionReactor.drainThrough

Relevant code:

  • OpenCodeAdapter.ts: cancelPendingOpenCodePrompt awaits Fiber.interrupt(admission.promptFiber) and Deferred.await(admission.submissionSettled) without a timeout, before the bounded remote-abort operations.
  • ThreadDeletionReactor.ts: drainThrough awaits the shared worker.drain after the event watermark.
  • ws.ts: explicit thread creation and first-message bootstrap both await this drain.

The traces identify the cancellation subtree but do not distinguish whether the stuck await was fiber interruption or submissionSettled.

Workaround

Restarting the headless T3 service after confirming no active AI turns cleared the stall. The same accepted-command replay then completed in 18 ms. HTTP-only health checks missed the failure.

There was also a second server started by npx t3 status during troubleshooting. It started after the first observed timeout; stopping that duplicate alone did not clear the stall. Restarting the primary process did.

Suggested regression coverage

Hold an OpenCode prompt cancellation open, delete its thread, then create/start an unrelated thread. Verify teardown terminates or reports a bounded failure and unrelated thread creation still completes. Include cancellation before prompt-fiber assignment and during submission, so completion is guaranteed across both paths.

Activity

  1. juliusmarminge commented on Sep 15, 2026

    @juliusmarminge
    Member

    Triage

    Confirmed. This is a real server availability bug in current main, not a one-off 0.0.40 glitch. A stuck OpenCode thread deletion can block every new chat on that process.

    What we see

    The reported sequence matches the code and the live probe:

    1. Delete an OpenCode thread while prompt admission/cancellation is still unsettled.
    2. processThreadDeleted stays in cancelPendingOpenCodePrompt (observed 242,562,870 ms, then “Interrupted during server shutdown”).
    3. A later thread.create / first-message bootstrap persists the empty thread, then waits in ThreadDeletionReactor.drainThrough and never reaches thread.turn-start-requested.
    4. The client remains connected; HTTP still succeeds. Restart clears the stall (replay acknowledged in 18 ms).

    The exact admission race that left cancellation unsettled is not deterministic. The global drain block is.

    Why

    Two cooperating defects.

    1. OpenCode teardown waits without a bound

    cancelPendingOpenCodePrompt does:

    • Fiber.interrupt(admission.promptFiber) when the fiber exists
    • then Deferred.await(admission.submissionSettled)

    Neither has a timeout. Remote abort (abortOpenCodeSessionForTeardown, 1s) runs after this wait, so a stuck cancel never reaches the already-bounded abort. The same unbounded pair is in interruptTurn.

    submissionSettled is only completed by the early-cancel check in sendTurn, or by promptEffect's onExit. Admission is installed before promptFiber is assigned, so a delete in that window waits on a deferred that sendTurn has not completed yet.

    The in-flight path is worse. runOpenCodeSdk is Effect.tryPromise. Interrupt aborts the AbortSignal, but the fiber does not finish until the promise settles. If session.promptAsync ignores abort, Fiber.interrupt never returns, onExit never runs, and submissionSettled never completes. The 10s promptAsync timeout is part of that same fiber: interrupting it also cancels the timeout, so the bound disappears exactly when teardown needs it.

    Existing coverage bounds a hung abort (aborts a held teardown request before closing the session scope). It does not hold prompt cancellation open across stopSession.

    2. One stuck cleanup is a process-wide create barrier

    ThreadDeletionReactor uses a single DrainableWorker. drainThrough waits for the event watermark, then worker.drain (queue empty and current item finished). ws.ts awaits that on explicit thread.create and on first-message bootstrap, so a reused thread id cannot own provider/terminal resources before the prior incarnation’s cleanup finishes (#8226).

    That fence is global, not per-thread. One hung providerService.stopSession keeps outstanding > 0, so every later create waits. Failures are logged and skipped; a hang is neither, so the worker never idles. Shutdown interrupting the 67-hour span matches a stuck worker item, not a crashed process.

    Not a duplicate

    Related but distinct:

    • #5241 — orphaned opencode serve after backend crash
    • #9065 — bundled serve holding OpenCode SQLite so CLI run hangs
    • #11730 — OpenCode background delegations look Settled
    • #8796 — forget the deleted thread’s provider binding; still calls unbounded stopSession on the same shared worker
    • #11613 / #8939 / #10805 — other OpenCode admission/abort work; none bound this wait or isolate drainThrough

    Next step

    Keep this open. Fix both layers; either one alone still leaves an instance-level stall.

    1. Bound OpenCode cancel/teardown. cancelPendingOpenCodePrompt and the matching interruptTurn wait should finish or fail in a short, explicit period (timeout Fiber.interrupt / Deferred.await, or Fiber.interruptFork plus a deferred timeout), then still run the existing remote abort. Do not let a non-settling SDK promise pin stopSession.
    2. Stop using one in-flight deletion as a process-wide create barrier. drainThrough should wait only for deletions that can conflict with the new thread (same id / reincarnation), or fail open after a bound so unrelated creates proceed. A hung cleanup can stay isolated and logged.
    3. Regression. Hold an OpenCode prompt cancel open (before promptFiber assignment and during an uninterruptible/hanging promptAsync), delete that thread, then create/start an unrelated thread. Teardown must terminate or report a bounded failure; the new thread must complete without restart.

    Workaround remains: confirm no active turns, restart the headless T3 process. HTTP health checks will not see this.

  2. added
    bugSomething is broken or behaving incorrectly.
    acceptedfeature request accepted
    via-triageFiled through npx t3 triage
    on Sep 15, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions