Skip to content

[Bug]: Stopped Claude thread silently loses native context when switching compatible provider instances #4766

Description

@reed-yang

Update 2026-09-02: reproduced on v0.0.38; see the latest comment and proposed fix #6148.

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. Configure two enabled Claude provider instances, A and B, that use the
    claudeAgent driver and resolve to the same Claude home / continuation key.
  2. Start a T3 thread on instance A and complete enough turns to persist a
    non-empty Claude resume cursor.
  3. Leave the thread idle until its provider session is soft-stopped.
  4. In the existing thread, select instance B before sending the next message.
  5. Send a short reply that depends on the immediately preceding assistant
    message, such as A in response to a list of proposed options.

The failure boundary is important: switching compatible instances while the
old provider session is active has a path that carries
activeSession.resumeCursor; switching after the session has stopped does
not.

Expected behavior

Because A and B advertise the same continuation identity, T3 should explicitly
carry the prior Claude resume cursor into B. If that is not safe for a given
pair of instances, T3 should reject the switch or require a transcript
handoff before sending the message.

A non-empty T3 thread should never silently start a blank provider-native
session.

Actual behavior

T3 keeps the full message history visible in the same UI thread but starts a
new Claude Agent SDK session without a resume cursor. Claude receives only the
new short message and replies as if the earlier conversation never happened.

In the observed incident:

  • The old Claude JSONL remained present and readable.
  • Neither the old nor new JSONL contained a compact_boundary entry.
  • The execution host, cwd, model, Claude home, driver, and continuation key
    were unchanged.
  • The replacement native session's persisted cursor became turnCount: 1.
  • T3's projected messages and checkpoints remained intact.

This rules out context-window exhaustion, compaction, missing native session
data, or a UI-history deletion.

The stopped-session path appears to lose the previous instance identity before
the turn starts:

  1. The model-selection update to B is projected before the turn-start event.
  2. With no active session, ProviderCommandReactor.ensureSessionForThread
    derives currentInstanceId from the already-updated
    thread.modelSelection.instanceId.
  3. The reactor therefore sees B as both the current and desired instance and
    starts it without explicitly carrying A's cursor.
  4. ProviderService.startSession only falls back to the persisted cursor when
    persistedBinding.providerInstanceId === resolvedInstanceId.
  5. The persisted binding belongs to A while the target is B, so the cursor is
    rejected even though both instances have the same continuation key.
  6. ClaudeAdapter consequently supplies a newly generated sessionId rather
    than resume to the Agent SDK.

Related but distinct reports:

Impact

Major degradation or frequent failure

The reset is silent and the UI still displays the old conversation, so the
user cannot tell that approvals, architectural decisions, constraints, and
prior tool results are absent from the model's context.

Version or commit

Observed on v0.0.29 (1153afb4). The relevant logic is unchanged on main at
887dd6e455bb / v0.0.30-nightly.20260728.933.

Environment

T3 Code Desktop on macOS with a remote Linux execution host; Claude Code
2.1.218 via two configured claudeAgent instances; same model, host, cwd,
Claude home, and continuation key across the switch.

Logs or stack traces

# Last context-bearing turn completes on instance A.
# The provider session is later soft-stopped while idle.
# The user selects compatible instance B immediately before the next turn.

provider.instance_id=B
provider.resume_cursor.present=false
provider.resume_cursor.source=none
claude.resume.source=generated-session
claude.query.resume=""

# After the context-free response:
resumeCursor.turnCount=1

No credentials, prompt contents, project paths, or full local identifiers are
included above.

Suggested fix and regression test

Resolve the previous provider instance and resume cursor from the persisted
binding or stopped session before the newly projected model selection replaces
that identity. Then:

  • If continuation identities match, explicitly pass the prior cursor to the
    target instance.
  • If they do not match, reject the switch or use a reviewed transcript handoff.
  • Emit a warning/activity entry when a thread with completed turns starts a
    provider session without a resume cursor.

Regression test:

  1. Start a Claude thread on instance A.
  2. Persist a non-empty cursor and soft-stop the session.
  3. Update the selection to compatible Claude instance B.
  4. Start the next turn.
  5. Assert that B receives A's prior cursor, or that the switch is explicitly
    rejected before the user message is sent. It must not silently start blank.

Screenshots, recordings, or supporting files

No response.

Workaround

Use one provider instance per T3 thread. When changing route or provider
instance, start a new T3 thread with a reviewed handoff. If the original
Claude session still exists, resume it directly from the same Claude home and
working directory.

Activity

  1. reed-yang commented on Sep 2, 2026

    @reed-yang
    Author

    Reproduced again on T3 Code v0.0.38 with the current claudeAgent adapter.

    This was a same-thread switch between two enabled Claude provider instances that use the same driver, Claude home, working directory, and continuation identity. The instances differ only in account/upstream routing.

    Sanitized trace sequence:

    t0  stopped thread is bound to instance A with a non-empty Claude resume cursor
    
    t1  first continuation on A:
        provider.resume_cursor.source=persisted
        provider.resume_cursor.present=true
        claude.resume.source=resume-session
    
    t2  the turn is interrupted and the session becomes stopped
    
    t3  model/provider selection is updated from A to compatible instance B
    
    t4  next turn starts on B:
        provider.resume_cursor.source=none
        provider.resume_cursor.present=false
        claude.resume.source=generated-session
        claude.query.resume=""
        claude.query.session_id=<new UUID>
    

    The CLI was therefore launched with --session-id <new UUID>, not --resume <old UUID>. The following turn correctly used --resume, but it resumed the newly created blank session rather than the original conversation.

    The T3 UI continued showing the complete projected thread history and displayed no continuity warning. Because T3 sends only the new prompt and relies on the provider-native resume cursor for historical context, Claude received none of the visible prior conversation.

    This is particularly dangerous for long-running coding work: the new agent can continue modifying the same repository while incorrectly believing it has the earlier decisions, constraints, approvals, and verification evidence.

    Current main still appears to contain the failing condition in ProviderService.startSession: persisted state is reused only when persistedBinding.providerInstanceId === resolvedInstanceId. In the stopped-session path, ProviderCommandReactor can also derive currentInstanceId from the already-updated model selection, so the source instance and its cursor are no longer available when the target session starts.

    PR #6148 appears to implement the correct fix by comparing continuation identities instead of instance IDs and carrying the persisted cursor/cwd across compatible Claude instances.

    Suggested regression coverage:

    1. Persist a stopped Claude session on instance A with a non-empty resume cursor.
    2. Update the requested model selection to compatible instance B before the next turn.
    3. Start the next turn through ProviderCommandReactor.
    4. Assert that B receives A's exact resume cursor.
    5. Assert that no generated session ID is created and the old binding is not overwritten.
    6. For incompatible instances, fail visibly before adapter startup.
    7. Never permit a non-empty projected T3 thread to silently start a blank provider-native session.

    This is a fresh real-world reproduction of #4766 on the current stable release.

  2. seido-agent commented on Sep 22, 2026

    @seido-agent

    Same root cause on the codex driver in v0.0.42, with a worse failure mode.

    Setup: two codex instances sharing CODEX_HOME, same cwd and continuation key, so the existing driver/continuation guards pass.

    Sequence:

    t0  stopped thread, binding on instance X, resume cursor set
    t1  switch to instance Y in the model picker, send a message
    t2  startSession: persisted cursor is dropped because the binding still names X
          effectiveResumeCursor = input.resumeCursor ?? (persistedBinding.providerInstanceId === resolvedInstanceId ? ... : undefined)
    t3  openCodexThread: resumeThreadId === undefined  ->  thread/start (never thread/resume)
    t4  a new, empty Codex thread is created; the UI still shows the full transcript
    

    Confirmed on disk: the original rollout is untouched and a new one appears.

    Why it is worse than the Claude case: a blank Codex thread still loads the repo files and working directory. With no history, the agent re-derives a role from them and answers confidently as a different agent than the thread it appears in. Nothing errors. For unattended setups in full-access mode that is a correctness risk, not just lost context.

    On #6148: the trigger here is the no-cursor path, not a failed resume — resumeThreadId === undefined never attempts thread/resume, so surfacing resume failures would not catch it. Worth carrying the cursor at startSession when continuation keys match, and refusing thread/start for a thread whose binding has ever held a cursor.

    Workaround: point the binding at the target instance before sending, so the cursor is inherited:

    UPDATE provider_session_runtime SET provider_instance_id = '<target-instance>' WHERE thread_id = '<thread-id>';

    Then check that the provider's existing rollout file grew instead of a new one appearing.

  3. YuriiHoliuk commented on Sep 24, 2026

    @YuriiHoliuk

    Still reproduces on 0.0.43-nightly.20260922.2110 (macOS 26.6, Claude Code 2.1.280).

    One thing the earlier reports don't cover: with more than one Claude subscription, usage limits make this the normal way to switch accounts, not an edge case.

    My setup is three claudeAgent instances, one per subscription, plus a fourth that picks an account at launch. All four use the same homePath (~/.claude), so they share a continuation key. Each one runs the stock claude binary with a different CLAUDE_CODE_OAUTH_TOKEN.

    This is what happens when an account runs out:

    1. The turn fails with "Claude usage limit reached". The session stays alive, in error.
    2. I wait for the reset or go do something else. After 30 idle minutes (DEFAULT_INACTIVITY_THRESHOLD_MS, checked every 5 minutes) the reaper stops the session.
    3. I pick another account in the model picker and send "continue". The session is stopped now, so the reactor's live-switch path that carries activeSession.resumeCursor doesn't run. ProviderService.startSession then drops the persisted cursor because persistedBinding.providerInstanceId !== resolvedInstanceId.

    The limit is the reason to switch, and waiting on it is what gets the session reaped. I rarely switch within 30 minutes, so for me this path fires nearly every time.

    It happened three times on two threads in one evening. The provider event logs (logs/provider/events.<thread>.log) show the split cleanly. Times are UTC, instances anonymized:

    thread 1
      16:30:24  session.started  A -> B while the session was live   payload.resume.resume = <original>   context kept
      17:49     usage limit on B
      18:24     session.exited (reaper)
      18:40:40  session.started  B -> C   payload = {}   blank
      18:41:13  session.started  C -> B   payload = {}   blank
    
    thread 2
      15:43     usage limit on C
      16:24     session.exited (reaper)
      16:30:05  session.started  C -> B   payload = {}   blank
    

    The first blank session on thread 2 did more damage than an error would have. The new agent got "continue", found the previous session's JSONL under ~/.claude/projects/, read parts of it, and worked from what it pieced together for 75 minutes. That included merging a PR. The UI still showed the whole conversation, so nothing looked wrong. I only found out later, when the next switch gave me an agent that said outright it had no earlier context.

    Recovering after the fact takes more than the provider_instance_id update suggested above. By the time you notice, the blank session has replaced resume_cursor_json with its own cursor (turnCount: 1), so pointing the binding at a new instance just resumes the blank session. If the provider log hasn't rotated away yet, the original cursor is in the session.started event from the last good start:

    grep -h '"type":"session.started"' ~/.t3/userdata/logs/provider/events.<thread-id>.log* \
      | grep <original-session-id>

    That cursor's resumeSessionAt is from when that session started, and the SDK truncates the resume at that message. Before writing it back, change resumeSessionAt to the uuid of the last top-level assistant message in the original JSONL. Then set resume_cursor_json to the edited cursor and provider_instance_id to the instance you're about to send on, while the thread is stopped. That brought back both of my threads.

    +1 for #6148. The warning the issue suggests, for a thread with completed turns starting a session without a resume cursor, would have caught all three of these.

  4. Nintorac commented on Oct 7, 2026

    @Nintorac

    FWIW I've hit this from a different direction, in my setup I'm running t3 in a container, the transcripts aren't mapped. I have a script to restore the transcripts from where they're stored and in older version of t3code I remembered to run the script by getting a warning telling me that the rollout was missing.

    So another test scenario might be,

    1. start a conversation,
    2. delete the transcript from under it
    3. resume the conversation

    here we should expect a warning/failure, might be good to go on step further and store a hash the transcript and compare expected vs actual

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions