Skip to content

Pi: resumed thread runs its turn with an invalidated extension ctx — T3's own pi-t3-mcp-extension hooks throw 'stale after session replacement' on every provider call #12467

Description

@astarktc

What happened

Resumed a day-old Pi thread in T3 Code (work Mac, Orchestrator V2 branch build) and sent a message. The turn ran, but every Pi extension hook that touched its ctx threw This extension ctx is stale after session replacement or reload…, rendered as red rows in the thread:

  • cmux-session failed during before_agent_start (1×)
  • index failed during before_agent_start (1×, a third-party extension)
  • pi-t3-mcp-extension failed during before_provider_request (3× — once per provider call in the turn)

The last one is T3's own injected extension (piT3McpExtensionSource.ts). Its before_provider_request handler reads only the hook-passed ctx (ctx.model?.provider) and captures nothing across sessions, so the ctx handed to the handler was already invalidated — this is a dispatch-side problem, not an extension holding a stale reference. Every T3-owned hook on that turn silently no-op'd: the OpenRouter max_tokens cap and, on other paths, the tool_call permission gate.

Resuming the same Pi session from a terminal (pi --session <file>) and sending the same message: no errors.

Diagnosis

T3 resumes a Pi thread by spawning pi --mode rpc (no --session), which opens a default session and loads all extensions against it, then sends switch_session to the persisted file (PiAdapterV2.ts registerThread, existing.nativeThreadRef.nativeId). In Pi 0.85.1 that switch is a full session replacement: teardownCurrent() → session.dispose() → extensionRunner.invalidate(...), then createRuntime() builds a fresh resource loader, re-runs every extension factory and installs a new runner. Any hook dispatched through the old runner after that point gets an invalidated ctx and throws on first access.

Driving this exact sequence by hand against pi --mode rpc (spawn → switch_session → prompt; also with the full global extension set; also after the default session had already run a turn) does not reproduce — the replaced runner is fresh and every hook is fine. So the plain switch is sound; the failing process had T3-specific history (kept alive across a day; whatever RPC sequence T3 issued to it between yesterday's last turn and today's resume) that left the dispatching runner invalidated while a turn still ran through it. I could not extract that sequence: the server trace does not record Pi RPC commands or the Pi child's stderr, and extension_error rows are not persisted in the event store, so the evidence is the UI only.

Two observations from the probes that may help whoever knows the pool lifecycle: (1) session_start with reason: "resume" fires twice per switch_session in RPC mode — rpc-mode.ts rebinds once via runtimeHost.setRebindSession inside finishSessionReplacement and again in case "switch_session", and each bindExtensions emits session_start; harmless in isolation, but it means every T3 resume double-runs every extension's session_start. (2) The whole class exists only because T3 replaces a session inside a live process. Spawning the resume as pi --mode rpc --session <file> (Pi's native resume; --session accepts the file path T3 already stores as nativeThreadRef.nativeId) would load extensions once against the right session and remove switch_session from the resume path entirely — no runner invalidation can then reach a turn. Every other CLI harness T3 drives is resumed by passing the session id on the command line; Pi supports the same.

Steps to reproduce

Not deterministic yet. Observed twice (2026-09-10 and 2026-09-18) with the same shape: a Pi thread idle for ≥ 1 day, reopened in T3, first message sent. A raw-RPC replay of spawn → switch_session → prompt does not trigger it (see Diagnosis).

Version

0.0.42 — Orchestrator V2 branch (t3code/codex-turn-mapping), head 934da5749, packaged desktop build.

Environment

macOS 15 (arm64), Node 24, Pi @earendil-works/pi-coding-agent 0.85.1 on both machines (one reproduced, one is the probe machine).

Evidence

UI rows (verbatim prefix; full text is Pi's standard stale-ctx message):

cmux-session failed during before_agent_start. This extension ctx is stale after session replacement or r…
index failed during before_agent_start. This extension ctx is stale after session replacement or reload. D…
pi-t3-mcp-extension failed during before_provider_request. This extension ctx is stale after session repl…
pi-t3-mcp-extension failed during before_provider_request. This extension ctx is stale after session repl…
pi-t3-mcp-extension failed during before_provider_request. This extension ctx is stale after session repl…

Pi source (0.85.1) for the replacement path: dist/core/agent-session-runtime.js switchSession() → teardownCurrent() → dist/core/agent-session.js dispose() line 595 this._extensionRunner.invalidate(...); dist/modes/rpc/rpc-mode.js case "switch_session" → rebindSession().

Clean probe transcript (spawn → switch → prompt, probe extension logging the session file from the hook ctx):

[probe] factory run ulnjx pid=77202
[probe:ulnjx] session_start reason=startup … 2026-09-18T15-34-03 (default session)
[probe] factory run de71e pid=77202          ← factories re-run on switch
[probe:de71e] session_start reason=resume … 2026-09-18T15-33-34 (target)
[probe:de71e] session_start reason=resume … (fired twice)
{"id":"2","type":"response","command":"switch_session","success":true,"data":{"cancelled":false}}
[probe:de71e] before_agent_start ok
[probe:de71e] before_provider_request ok

Related issues

None found for stale ctx / switch_session. #12285 (delegated-wake cap) and #11168 (mode:"wait") are different V2 contract problems from the same setup; not duplicates.

Fix applied or workaround

User-side: every hand-built extension here now wraps pi.on with a guard that logs the stale throw once and no-ops afterwards (that is why cmux-session shows 1 row instead of one per call). That hides the symptom for our extensions only; T3's own extension and third-party ones still throw, and the guarded hooks still do nothing for the turn. Reliable recovery is to resume the session from a terminal instead of T3.

Filed by

claude (opus-5) via a Pi thread inside T3 Code, following the t3 triage report structure.

Activity

  1. juliusmarminge commented on Sep 18, 2026

    @juliusmarminge
    Member

    Triage

    Confirmed as a real Orchestrator V2 Pi resume bug. The report matches the current V2 branch (t3code/codex-turn-mapping / PR #2829, HEAD 934da5749, still 0.0.42). This code is not on main. Not a duplicate of #12285 or #11168 — those are different V2 MCP contract problems from the same setup.

    What breaks

    T3 never resumes Pi the way Pi itself does. PiAdapterV2.openSession always spawns pi --mode rpc with no --session (buildPiRpcLaunch in piT3McpInjection.ts). registerThread / resumeThread then send switch_session to existing.nativeThreadRef.nativeId (the session file path). In Pi 0.85.1 that switch is a full session replacement: teardownCurrent() → session.dispose() → extensionRunner.invalidate(...), then factories re-run against a new runner. Any hook dispatched through the old runner throws on first ctx access.

    That is exactly the UI fingerprint:

    • cmux-session / index during before_agent_start (user extensions that touch ctx)
    • pi-t3-mcp-extension during before_provider_request (once per provider call)

    T3’s own extension (piT3McpExtensionSource.ts) is not holding a stale closure. before_provider_request only reads the hook-passed ctx.model?.provider. The ctx handed to the handler is already invalidated — dispatch-side.

    On that turn the T3 hooks that do touch ctx silently no-op:

    • OpenRouter max_tokens cap (before_provider_request)
    • Supervised / auto-accept permission gate (tool_call → ctx.ui.confirm)

    T3’s before_agent_start orchestration-instruction injection does not read ctx, so it can still apply. The turn still runs. Recovery from a terminal (pi --session <file>) is clean because that path never replaces a live runner.

    What the day-old idle actually is

    ProviderSessionManager idle-releases after 30 minutes (max pin 4 hours). Pi does not implement hasPendingBackgroundWork, so a ≥1 day idle thread is a fresh openSession, not a day-old RPC process. The “kept alive across a day” history is the persisted T3 thread / nativeThreadRef, not the child.

    Fresh resume sequence on this branch:

    1. Spawn pi --mode rpc (default session; every extension factory runs against it).
    2. Fork get_commands immediately — the adapter comment says discovery “can invoke extension code” and it is not awaited before switch_session.
    3. resumeThread → switch_session → runner invalidate + factory re-run. session_start reason=resume fires twice per switch in Pi RPC (rebind in finishSessionReplacement and again in case "switch_session"). Harmless alone; it does mean every T3 resume double-runs session_start.
    4. get_state / get_entries, then startTurn → set_model / set_session_name / prompt.

    A hand replay of spawn → switch_session → prompt does not reproduce. The T3-specific delta is that overlapping get_commands plus the rest of the register/turn RPC on a process that just invalidated its runner. Server traces do not record Pi RPC command types; stderr is length-only; extension_error is a UI turn item, not an event-store row. That is why the evidence is the UI only.

    Why this is T3’s resume shape, not “Pi switch is broken”

    The V2 open-session contract already has the native file path at spawn time:

    • ProviderAdapterV2OpenSessionInput.initialNativeThreadId — “Native thread to activate while an eager adapter opens its provider process.”
    • ProviderTurnStartService already passes providerThread.nativeThreadRef.nativeId (the session file).
    • ACP uses that field for spawn-time resume. Pi ignores it and always opens a default session, then switches.

    Pi’s CLI already accepts this: pi --mode rpc --session <path|id>. T3 already stores that path as nativeThreadRef.nativeId. Every other CLI harness T3 drives is resumed by passing the session id at start. User launch-arg --session should stay rejected — that reservation is “T3 owns session identity,” not “T3 must not pass --session itself.”

    switch_session is still required for in-process fork adoption (forkThread creates a new file, then registerThread on the live process). It should leave the cold-resume path.

    Related

    Suggested fix (on the V2 branch)

    1. Spawn the resume as pi --mode rpc --session <initialNativeThreadId> when openSession receives that field. Keep rejecting --session / --session-id / --resume from user launchArgs.
    2. Skip switch_session when get_state.sessionFile already matches the persisted native id. A same-file switch still tears down and replaces the runner.
    3. Keep switch_session for fork adoption on a live process.
    4. Optional, for the next bug in this class: log Pi RPC command types (not payloads) around session bind; don’t start get_commands until the thread is bound.

    Existing adapter test (“registers the thread from get_state and resumes via switch_session”) encodes the current cold-resume shape and will need to expect spawn-time --session plus no switch when the file already matches.

    No change needed on main. Fix belongs on t3code/codex-turn-mapping.

    Workaround until then: resume the same session from a terminal with pi --session <file>. Guarding pi.on in user extensions only hides their rows; T3’s hooks still no-op.

  2. added
    bugSomething is broken or behaving incorrectly.
    acceptedfeature request accepted
    via-triageFiled through npx t3 triage
    on Sep 18, 2026
  3. added a commit that references this issue on Oct 2, 2026
    66af5b1
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions