Skip to content

Interrupt timeouts report a terminal status for work that may still be running #16844

Description

@Adamulek123

Problem

Three adapters settle a turn as terminal even when the native run may still be alive, after a fixed timeout. The turn is reported finished or interrupted while the provider process is still working, and the app stops ingesting from it.

  • CursorAdapterV2.ts:2458-2469 — context.run.cancel is issued at 2457, then after Deferred.await(context.completed).pipe(Effect.timeoutOption("10 seconds")) the adapter calls finalizeTurn. Nothing verifies the native agent actually stopped.
  • ClaudeAdapterV2.ts:7330-7351 — the same 10s bound, then finalizeActiveTurn, plus Deferred.succeed(existing.closed, undefined). That Deferred.succeed races the real stream's own Effect.ensuring(Deferred.succeed(closed, undefined)) at 7096, so the adapter's own exit path can observe a closed query that is still streaming.
  • CodexAdapterV2.ts:1827-1838 awaitActiveTurn — a 1000-iteration Effect.yieldNow spin that returns undefined on exhaustion. Every caller then either drops the delta (2933, 3841, 4213) or calls toProtocolError (4476, 4539, 4600, 4865), which the app-server reads as a denial.

Why this is a design question, not a bug

Should a 10-second bound be allowed to lie about terminal state at all? Two defensible answers:

  1. Yes — a stuck provider must not hold the thread open forever, so time-box it and report honestly that we gave up.
  2. No — report interrupted/unknown rather than a terminal success, and keep a watchdog that escalates.

Option 1 is what the code does today. The problem is that it reports a terminal success-shaped status for work that may not have finished, and the user cannot tell the difference from a real completion.

Impact

#15197 is the user-visible symptom: a V2 Claude run stays "running" forever after the CLI reports result/success. #15542 is adjacent but a different cause (a failed commit that loses supervision, rather than a timeout that lies on purpose).

The 10s bound pattern was introduced by v1 PR #7349, which was closed unmerged.

Environment

Found by static audit of the merged orchestrator-v2 work (#2829). Re-verified present at main. #15778 already deferred the related "maintainer lifecycle decision" for approval survival and interrupt, so the routing is established.

Deliberately not filed as a PR: these are five files and four providers with genuinely different risk profiles, and the fix depends on the answer to the question above.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions