Repository navigation
[Bug]: Stuck OpenCode deletion blocks new chats in ThreadDeletionReactor.drainThrough #11889
Description
Activity
Triage
Confirmed. This is a real server availability bug in current
main, not a one-off 0.0.40 glitch. A stuck OpenCode thread deletion can block every new chat on that process.What we see
The reported sequence matches the code and the live probe:
- Delete an OpenCode thread while prompt admission/cancellation is still unsettled.
processThreadDeletedstays incancelPendingOpenCodePrompt(observed 242,562,870 ms, then “Interrupted during server shutdown”).- A later
thread.create/ first-message bootstrap persists the empty thread, then waits inThreadDeletionReactor.drainThroughand never reachesthread.turn-start-requested. - The client remains connected; HTTP still succeeds. Restart clears the stall (replay acknowledged in 18 ms).
The exact admission race that left cancellation unsettled is not deterministic. The global drain block is.
Why
Two cooperating defects.
1. OpenCode teardown waits without a bound
cancelPendingOpenCodePromptdoes:Fiber.interrupt(admission.promptFiber)when the fiber exists- then
Deferred.await(admission.submissionSettled)
Neither has a timeout. Remote abort (
abortOpenCodeSessionForTeardown, 1s) runs after this wait, so a stuck cancel never reaches the already-bounded abort. The same unbounded pair is ininterruptTurn.submissionSettledis only completed by the early-cancel check insendTurn, or bypromptEffect'sonExit. Admission is installed beforepromptFiberis assigned, so a delete in that window waits on a deferred thatsendTurnhas not completed yet.The in-flight path is worse.
runOpenCodeSdkisEffect.tryPromise. Interrupt aborts theAbortSignal, but the fiber does not finish until the promise settles. Ifsession.promptAsyncignores abort,Fiber.interruptnever returns,onExitnever runs, andsubmissionSettlednever completes. The 10spromptAsynctimeout is part of that same fiber: interrupting it also cancels the timeout, so the bound disappears exactly when teardown needs it.Existing coverage bounds a hung abort (
aborts a held teardown request before closing the session scope). It does not hold prompt cancellation open acrossstopSession.2. One stuck cleanup is a process-wide create barrier
ThreadDeletionReactoruses a singleDrainableWorker.drainThroughwaits for the event watermark, thenworker.drain(queue empty and current item finished).ws.tsawaits that on explicitthread.createand on first-message bootstrap, so a reused thread id cannot own provider/terminal resources before the prior incarnation’s cleanup finishes (#8226).That fence is global, not per-thread. One hung
providerService.stopSessionkeepsoutstanding > 0, so every later create waits. Failures are logged and skipped; a hang is neither, so the worker never idles. Shutdown interrupting the 67-hour span matches a stuck worker item, not a crashed process.Not a duplicate
Related but distinct:
- #5241 — orphaned
opencode serveafter backend crash - #9065 — bundled serve holding OpenCode SQLite so CLI
runhangs - #11730 — OpenCode background delegations look Settled
- #8796 — forget the deleted thread’s provider binding; still calls unbounded
stopSessionon the same shared worker - #11613 / #8939 / #10805 — other OpenCode admission/abort work; none bound this wait or isolate
drainThrough
Next step
Keep this open. Fix both layers; either one alone still leaves an instance-level stall.
- Bound OpenCode cancel/teardown.
cancelPendingOpenCodePromptand the matchinginterruptTurnwait should finish or fail in a short, explicit period (timeoutFiber.interrupt/Deferred.await, orFiber.interruptForkplus a deferred timeout), then still run the existing remote abort. Do not let a non-settling SDK promise pinstopSession. - Stop using one in-flight deletion as a process-wide create barrier.
drainThroughshould wait only for deletions that can conflict with the new thread (same id / reincarnation), or fail open after a bound so unrelated creates proceed. A hung cleanup can stay isolated and logged. - Regression. Hold an OpenCode prompt cancel open (before
promptFiberassignment and during an uninterruptible/hangingpromptAsync), delete that thread, then create/start an unrelated thread. Teardown must terminate or report a bounded failure; the new thread must complete without restart.
Workaround remains: confirm no active turns, restart the headless T3 process. HTTP health checks will not see this.
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.acceptedfeature request acceptedfeature request acceptedvia-triageFiled through npx t3 triageFiled through npx t3 triage
on Sep 15, 2026
Before submitting
Area
apps/server
Summary
Deleting an OpenCode thread left
processThreadDeletedstuck incancelPendingOpenCodePromptfor 242,562,870 ms, about 67.4 hours. New thread creation then blocked inThreadDeletionReactor.drainThrough, although the remote environment remained connected and HTTP requests succeeded.Steps to reproduce
Observed sequence:
thread.createdevent is persisted, but the request waits inThreadDeletionReactor.drainThroughand never reachesthread.turn-start-requested.The exact timing that caused the original OpenCode cancellation to remain unsettled has not been reproduced deterministically. The blocked new-thread path was reproduced twice against the affected live process with the following harmless probe:
thread.createdcommand ID, payload, and receipt sequence from SQLite in read-only mode.thread.createcommand over authenticatedorchestration.dispatchCommand.This probe creates no additional thread and submits no model prompt. It timed out before recovery and passed in 18 ms after restarting the primary server.
Expected behavior
Provider teardown should finish or fail within a bounded period. A stuck deletion should not prevent unrelated new chats from starting.
Actual behavior
The client remains connected. New chat requests persist the empty thread, then wait indefinitely behind the shared deletion worker. Two original dispatch spans lasted 121,844.8 ms and 47,730.2 ms before client interruption. On service shutdown, the previously open deletion span finally appeared with a duration of 242,562,870 ms.
Impact
Blocks new chat creation on the affected server until recovery.
Version or commit
T3 server and desktop 0.0.40. The unbounded cancellation wait and shared drain are also present in the upstream main files inspected on September 15, 2026.
Environment
Linux x86_64 headless server, Node v22.23.2, macOS T3 desktop 0.0.40, OpenCode provider. Private paths, thread IDs, and chat content omitted.
Logs or stack traces
Sanitized summary of the shutdown trace:
Relevant code:
cancelPendingOpenCodePromptawaitsFiber.interrupt(admission.promptFiber)andDeferred.await(admission.submissionSettled)without a timeout, before the bounded remote-abort operations.drainThroughawaits the sharedworker.drainafter the event watermark.The traces identify the cancellation subtree but do not distinguish whether the stuck await was fiber interruption or
submissionSettled.Workaround
Restarting the headless T3 service after confirming no active AI turns cleared the stall. The same accepted-command replay then completed in 18 ms. HTTP-only health checks missed the failure.
There was also a second server started by
npx t3 statusduring troubleshooting. It started after the first observed timeout; stopping that duplicate alone did not clear the stall. Restarting the primary process did.Suggested regression coverage
Hold an OpenCode prompt cancellation open, delete its thread, then create/start an unrelated thread. Verify teardown terminates or reports a bounded failure and unrelated thread creation still completes. Include cancellation before prompt-fiber assignment and during submission, so completion is guaranteed across both paths.