Skip to content

[Bug]: Runtime-mode, instance and worktree changes detach the Claude session and kill live subagents without the background-work check #16000

Description

@OscarsOrbit

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. Start a Claude thread in full-access runtime mode.
  2. In one turn, start two pieces of background work and confirm both are progressing:
    • a background Bash command (run_in_background) that appends a counter line to a file every 5 s;
    • a background subagent (Agent tool, run_in_background) that runs 24 short timed steps (about 4 min) and appends one line per step to a second file.
  3. While both are still running, change the thread's runtime mode in the composer (Full access → Auto-accept edits) and send any message.
  4. Watch the thread, the two files, and the thread's orchestration events.

The worktree-change and provider-instance-switch paths go through the same detach (see Actual behavior); their trial results will follow in a comment.

Expected behavior

The same protection #14726 added for model and effort changes: while Claude still has background agents or commands running, a change that ends the Claude CLI process is refused or deferred with a clear, retryable message, and the background work is not lost. The guard's own comment gives the reason: "Background agents and shells run inside the CLI process, so a new selection would kill them and lose their results" (

// Background agents and shells run inside the CLI process, so a
// new selection would kill them and lose their results. Refuse until
// they finish or the user presses Stop, which closes the process.
// Another native thread on this session is one the app thread has
// left (Claude sessions serve one app thread), so it is replaced.
if (
existing !== null &&
existing.nativeThreadId === nativeThreadId &&
!existing.stopping &&
(yield* liveProcessRunsBackgroundWork(existing))
) {
return yield* new ClaudeBackgroundWorkBlocksQueryReplacementError();
}
).

Actual behavior

Sending the message with the new runtime mode ends the Claude session at once, with no refusal or warning, and the running subagent is killed. In one controlled trial (times UTC, from the thread's events and the two progress files):

  • 04:21:59.207 the message is sent: provider-session.detached, reason "Runtime mode changed."
  • 04:21:59.291 the subagent goes to failed, 84 ms later. Its progress file stopped at step 18 of 24 and its completion marker never appeared (checked at T+305 s).
  • 04:22:01.256 a new provider session attaches. The UI shows "Provider error — The provider event stream closed unexpectedly", the background agent as Failed mid-run, and the new session says the background command and subagent "were reported stopped when the previous session ended."

A note on background shells (not a confirmed T3 defect): in the same trial, our background Bash loop kept writing to its file after T3 and the new session reported it stopped (40 lines at the change, 64 at +120 s). We believe this is an artifact of our harness: the loop's timeout child is reparented and escapes the CLI teardown. We did not test a background command that stays inside the CLI process tree, so we report it only as an observation.

Source at efecd3c:

  • Runtime-mode, worktree and provider/instance changes go through dispatchThreadMutation, which emits provider-session.detached and a detach effect with no pending-background check (
    const detachSessionIds = new Set(
    command.type === "thread.archive" || command.type === "thread.settle"
    ? (providerContext?.providerSessions ?? []).map((session) => session.id)
    : command.type === "thread.metadata.update" &&
    command.worktreePath !== undefined &&
    command.worktreePath !== thread.worktreePath
    ? (providerContext?.providerSessions ?? []).map((session) => session.id)
    : command.type === "thread.runtime-mode.set"
    ? (providerContext?.providerSessions ?? [])
    .filter(
    (session) => !session.capabilities.sessions.supportsRuntimeModeSwitchInSession,
    )
    .map((session) => session.id)
    : (providerSwitchPlan?.releaseProviderSessionIds ?? []),
    );
    if (detachSessionIds.size > 0) {
    const liveSessions = (providerContext?.providerSessions ?? []).filter(
    (session) =>
    detachSessionIds.has(session.id) &&
    session.status !== "stopped" &&
    session.status !== "error",
    );
    yield* Effect.forEach(
    liveSessions,
    (session) =>
    Effect.gen(function* () {
    yield* emit(
    events,
    command,
    )({
    type: "provider-session.detached",
    threadId: command.threadId,
    driver: session.driver,
    providerInstanceId: session.providerInstanceId,
    occurredAt: now,
    payload: {
    providerSessionId: session.id,
    detachedAt: now,
    reason:
    command.type === "thread.archive"
    ? "Thread archived."
    : command.type === "thread.settle"
    ? "Thread settled."
    : command.type === "thread.metadata.update"
    ? "Workspace changed."
    : command.type === "thread.runtime-mode.set"
    ? "Runtime mode changed."
    : "Provider or model selection changed.",
    },
    });
    const pendingEffect = {
    id: `effect:${command.commandId}:provider-session.detach:${session.id}`,
    commandId: command.commandId,
    threadId: command.threadId,
    request: {
    type: "provider-session.detach",
    providerSessionId: session.id,
    detail:
    command.type === "thread.archive"
    ? "Thread archived."
    : command.type === "thread.settle"
    ? "Thread settled."
    : command.type === "thread.metadata.update"
    ? "Workspace changed."
    : command.type === "thread.runtime-mode.set"
    ? "Runtime mode changed."
    : "Provider or model selection changed.",
    // Terminal detaches revoke the thread's MCP credentials; other
    // detach reasons keep them so a re-attaching provider process
    // stays authorized.
    ...(command.type === "thread.archive" ? { revokeMcpCredential: true } : {}),
    },
    } satisfies PendingOrchestrationEffectV2;
    yield* Ref.update(effects, (existing) => [...existing, pendingEffect]);
    }),
    { concurrency: 1, discard: true },
    );
    }
    if (command.type === "thread.archive") {
    yield* Ref.update(effects, (existing) => [
    ...existing,
    ). Claude declares supportsRuntimeModeSwitchInSession: false (
    supportsRuntimeModeSwitchInSession: false,
    ).
  • An instance/runtime/workspace change returns restart_and_resume (
    const instanceChanged = current.modelSelection.instanceId !== target.modelSelection.instanceId;
    const runtimeChanged = current.runtimeMode !== target.runtimeMode;
    const workspaceChanged = current.workspace !== target.workspace;
    const selectionChanged = !modelSelectionsEqual(current.modelSelection, target.modelSelection);
    if (selectionChanged && !instanceChanged) {
    switch (input.selectionTransition?.type) {
    case "create_with_handoff":
    return { type: "create_with_handoff" };
    case "reject":
    return { type: "reject", reason: input.selectionTransition.reason };
    case undefined:
    return {
    type: "reject",
    reason: "The provider adapter did not classify the selection change.",
    };
    case "apply_on_next_turn":
    case "restart_session":
    break;
    }
    }
    if (instanceChanged || runtimeChanged || workspaceChanged) {
    return { type: "restart_and_resume" };
    }
    ); the switch plan releases the session (
    const releaseProviderSessionIds = projection.providerSessions
    .filter((session) => {
    if (!isLiveProviderSession(session)) return false;
    if (transition.type === "restart_and_resume") {
    // Other live records may serve pooled or delegated bindings;
    // only the session being replaced is released.
    return session.id === currentSession?.id;
    }
    if (transition.type === "create_with_handoff") {
    return session.providerInstanceId !== targetModelSelection.instanceId;
    }
    return false;
    })
    .map((session) => session.id);
    ).
  • Claude detach calls releaseEntry(reason: "manual_shutdown"), closing the CLI process (
    detach: (input) =>
    Effect.gen(function* () {
    const key = sessionKey(input.providerSessionId);
    const currentEntry = (yield* Ref.get(sessions)).get(key);
    let detachedProviderThreads: ReadonlyArray<OrchestrationV2ProviderThread> = [];
    if (currentEntry?.supportsMultipleProviderThreads === true) {
    const projection = yield* Effect.option(
    projectionStore.getThreadRecords(input.threadId, [
    "providerThreads",
    "providerTurns",
    ]),
    );
    if (Option.isSome(projection)) {
    const providerThreads = new Map(
    projection.value.providerThreads
    .filter((thread) => thread.providerSessionId === input.providerSessionId)
    .map((thread) => [thread.id, thread] as const),
    );
    detachedProviderThreads = [...providerThreads.values()];
    const activeTurns = projection.value.providerTurns.filter(
    (turn) => turn.status === "running" && providerThreads.has(turn.providerThreadId),
    );
    yield* Effect.forEach(
    activeTurns,
    (turn) =>
    currentEntry.exposedRuntime
    .interruptTurn({
    providerThread: providerThreads.get(turn.providerThreadId)!,
    providerTurnId: turn.id,
    })
    .pipe(
    Effect.catchCause((cause) =>
    Effect.logWarning(
    "orchestration-v2.driver-session.detach-interrupt-failed",
    {
    providerSessionId: input.providerSessionId,
    threadId: input.threadId,
    providerTurnId: turn.id,
    cause,
    },
    ),
    ),
    ),
    { concurrency: 1, discard: true },
    );
    }
    }
    const detached = yield* Ref.modify(sessions, (current) => {
    const entry = current.get(key);
    if (entry === undefined || !entry.attachedThreadIds.has(input.threadId)) {
    return [Option.none<LiveSessionEntry>(), current] as const;
    }
    const attachedThreadIds = new Set(entry.attachedThreadIds);
    attachedThreadIds.delete(input.threadId);
    const loadedProviderThreadKeyByThread = new Map(
    entry.loadedProviderThreadKeyByThread,
    );
    loadedProviderThreadKeyByThread.delete(input.threadId);
    // For a plain (workspace-change) detach, the credential id stays
    // recorded: the thread may re-attach and reuse it, and
    // releaseEntry revokes it when the provider process finally goes
    // away. A terminal detach (archive/delete) prunes the record so
    // nothing vetoes the revocation below.
    const mcpCredentialIdByThread =
    input.revokeMcpCredential === true
    ? (() => {
    const pruned = new Map(entry.mcpCredentialIdByThread);
    pruned.delete(input.threadId);
    return pruned;
    })()
    : entry.mcpCredentialIdByThread;
    const updatedEntry = {
    ...entry,
    attachedThreadIds,
    loadedProviderThreadKeyByThread,
    mcpCredentialIdByThread,
    };
    const updated = new Map(current);
    updated.set(key, updatedEntry);
    return [Option.some(updatedEntry), updated] as const;
    });
    // Plain detaches deliberately do not revoke: a detached thread's
    // provider process may still be alive (shared multi-thread codex
    // session across a workspace handoff) and holds its MCP client's
    // credential for the thread it will re-attach with. Credentials
    // are revoked when the session entry is released (process gone)
    // or rotated on the next attach if they stopped resolving.
    // Terminal detaches (thread archived or deleted) revoke the
    // thread's credentials immediately, even on a retry where the
    // entry is already gone: there is no legitimate future re-attach,
    // and the token must not outlive the thread.
    if (input.revokeMcpCredential === true) {
    yield* clearMcpSession(input.threadId);
    }
    if (Option.isNone(detached)) {
    return;
    }
    if (
    detached.value.attachedThreadIds.size === 0 &&
    !detached.value.supportsMultipleProviderThreads
    ) {
    yield* releaseEntry({
    providerSessionId: input.providerSessionId,
    reason: "manual_shutdown",
    ...(input.detail === undefined ? {} : { detail: input.detail }),
    ). hasPendingBackgroundWork is consulted only on idle release (
    const key = sessionKey(input.providerSessionId);
    const entry = current.get(key);
    if (
    entry === undefined ||
    entry.busyCount > 0 ||
    entry.idleGeneration !== input.generation
    ) {
    return;
    }
    // Capture runtime identity before yielding: a replacement session
    // can reuse the same providerSessionId while this fiber is parked.
    const probedRuntime = entry.runtime;
    const hasPendingWork =
    probedRuntime.hasPendingBackgroundWork === undefined
    ? false
    : yield* probedRuntime.hasPendingBackgroundWork.pipe(
    Effect.catchCause(() => Effect.succeed(false)),
    );
    if (hasPendingWork) {
    const now = yield* Clock.currentTimeMillis;
    const pinnedSinceMs = entry.pinnedSinceMs ?? now;
    if (now - pinnedSinceMs < maxIdlePinMs) {
    const shouldContinuePin = yield* Ref.modify(sessions, (latest) => {
    const latestEntry = latest.get(key);
    if (
    latestEntry === undefined ||
    latestEntry.busyCount > 0 ||
    latestEntry.idleGeneration !== input.generation ||
    latestEntry.runtime !== probedRuntime
    ) {
    return [false, latest] as const;
    }
    const updated = new Map(latest);
    updated.set(key, { ...latestEntry, pinnedSinceMs });
    return [true, updated] as const;
    });
    if (!shouldContinuePin) {
    // Generation or runtime advanced while we probed pending work;
    // the current owner of the entry owns idle release.
    return;
    }
    yield* Effect.logInfo("orchestration-v2.driver-session.idle-release-deferred", {
    providerSessionId: input.providerSessionId,
    pinnedForMs: now - pinnedSinceMs,
    });
    // Re-check on this fiber after another idle window. Do not call
    // scheduleIdleReleaseInternal: that cancels entry.idleFiber, which
    // is this fiber, and can self-deadlock on Fiber.interrupt.
    yield* Effect.sleep(Duration.millis(idleTimeoutMs));
    return yield* releaseIfStillIdle(input);
    }
    yield* Effect.logWarning("orchestration-v2.driver-session.idle-release-pin-expired", {
    providerSessionId: input.providerSessionId,
    pinnedForMs: now - pinnedSinceMs,
    });
    }
    // hasPendingBackgroundWork yields to the adapter, so the idle
    ); the fix(server): Claude model changes no longer kill running background agents #14726 guard covers query replacement only.

Searched before filing (open and closed): "runtime mode background subagent", "worktree handoff background", "kills background agents", "BackgroundWorkBlocks", "provider-session.detached" (only #15695, unrelated). The closest report is #12694, closed by #14726, which guards model and effort replacement; this is the adjacent unguarded set (runtime mode, instance, worktree).

Suggested fix: apply the pending-background-work check before the session detach/restart, as #14726 does for query replacement (refuse with the same retryable message, or defer until the background roster is empty); on a detach, stop or reattach tracked background shells rather than reporting them stopped.

Impact

Major degradation or frequent failure

Version or commit

v0.0.46-nightly.20261004.2657 (efecd3c)

Environment

Linux x64 container, self-hosted t3 server, Node 24.13.1, Claude Code 2.1.289, Codex 0.160.0; Claude via an API-compatible proxy (not involved)

Logs or stack traces

04:18:33.352Z provider-session.attached   status=ready
04:21:59.096Z subagent.updated            status=running   (last running update)
04:21:59.207Z provider-session.detached   reason="Runtime mode changed."
04:21:59.291Z subagent.updated            status=failed
04:22:01.256Z provider-session.attached   status=ready   (new session)
subagent progress file: 18 of 24 steps; no completion marker at T+305 s
background bash counter: 40 lines at the change, 64 at T+120 s, ran to its timeout (reported stopped)

Screenshots, recordings, or supporting files

No response

Workaround

Wait for background agents and commands to finish, or stop them, before changing the runtime mode, the provider instance or the worktree.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions