Repository navigation
Pi: resumed thread runs its turn with an invalidated extension ctx — T3's own pi-t3-mcp-extension hooks throw 'stale after session replacement' on every provider call #12467
Description
Activity
Triage
Confirmed as a real Orchestrator V2 Pi resume bug. The report matches the current V2 branch (
t3code/codex-turn-mapping/ PR #2829, HEAD934da5749, still 0.0.42). This code is not onmain. Not a duplicate of #12285 or #11168 — those are different V2 MCP contract problems from the same setup.What breaks
T3 never resumes Pi the way Pi itself does.
PiAdapterV2.openSessionalways spawnspi --mode rpcwith no--session(buildPiRpcLaunchinpiT3McpInjection.ts).registerThread/resumeThreadthen sendswitch_sessiontoexisting.nativeThreadRef.nativeId(the session file path). In Pi 0.85.1 that switch is a full session replacement:teardownCurrent()→session.dispose()→extensionRunner.invalidate(...), then factories re-run against a new runner. Any hook dispatched through the old runner throws on firstctxaccess.That is exactly the UI fingerprint:
cmux-session/indexduringbefore_agent_start(user extensions that touchctx)pi-t3-mcp-extensionduringbefore_provider_request(once per provider call)
T3’s own extension (
piT3McpExtensionSource.ts) is not holding a stale closure.before_provider_requestonly reads the hook-passedctx.model?.provider. The ctx handed to the handler is already invalidated — dispatch-side.On that turn the T3 hooks that do touch
ctxsilently no-op:- OpenRouter
max_tokenscap (before_provider_request) - Supervised / auto-accept permission gate (
tool_call→ctx.ui.confirm)
T3’s
before_agent_startorchestration-instruction injection does not readctx, so it can still apply. The turn still runs. Recovery from a terminal (pi --session <file>) is clean because that path never replaces a live runner.What the day-old idle actually is
ProviderSessionManageridle-releases after 30 minutes (max pin 4 hours). Pi does not implementhasPendingBackgroundWork, so a ≥1 day idle thread is a freshopenSession, not a day-old RPC process. The “kept alive across a day” history is the persisted T3 thread /nativeThreadRef, not the child.Fresh resume sequence on this branch:
- Spawn
pi --mode rpc(default session; every extension factory runs against it). - Fork
get_commandsimmediately — the adapter comment says discovery “can invoke extension code” and it is not awaited beforeswitch_session. resumeThread→switch_session→ runner invalidate + factory re-run.session_start reason=resumefires twice per switch in Pi RPC (rebind infinishSessionReplacementand again incase "switch_session"). Harmless alone; it does mean every T3 resume double-runssession_start.get_state/get_entries, thenstartTurn→set_model/set_session_name/prompt.
A hand replay of spawn →
switch_session→promptdoes not reproduce. The T3-specific delta is that overlappingget_commandsplus the rest of the register/turn RPC on a process that just invalidated its runner. Server traces do not record Pi RPC command types; stderr is length-only;extension_erroris a UI turn item, not an event-store row. That is why the evidence is the UI only.Why this is T3’s resume shape, not “Pi switch is broken”
The V2 open-session contract already has the native file path at spawn time:
ProviderAdapterV2OpenSessionInput.initialNativeThreadId— “Native thread to activate while an eager adapter opens its provider process.”ProviderTurnStartServicealready passesproviderThread.nativeThreadRef.nativeId(the session file).- ACP uses that field for spawn-time resume. Pi ignores it and always opens a default session, then switches.
Pi’s CLI already accepts this:
pi --mode rpc --session <path|id>. T3 already stores that path asnativeThreadRef.nativeId. Every other CLI harness T3 drives is resumed by passing the session id at start. User launch-arg--sessionshould stay rejected — that reservation is “T3 owns session identity,” not “T3 must not pass--sessionitself.”switch_sessionis still required for in-process fork adoption (forkThreadcreates a new file, thenregisterThreadon the live process). It should leave the cold-resume path.Related
- Not a duplicate of delegate_task async completions are silently starved after two wake deliveries per cohort — parent never wakes for third-wave children (V2 branch) #12285 (async
delegate_taskwake lifetime cap) or delegate_task mode:"wait" dies at the client ~5 min ceiling and returns no taskId, causing duplicate child dispatch #11168 (mode:"wait"~5 min client timeout, notaskId). - No prior issue for
stale ctx/switch_session. - Pi provider: feat(providers): add Pi coding agent #7211 on this branch. V2 umbrella: feat(orchestrator): introduce new orchestrator #2829.
Suggested fix (on the V2 branch)
- Spawn the resume as
pi --mode rpc --session <initialNativeThreadId>whenopenSessionreceives that field. Keep rejecting--session/--session-id/--resumefrom userlaunchArgs. - Skip
switch_sessionwhenget_state.sessionFilealready matches the persisted native id. A same-file switch still tears down and replaces the runner. - Keep
switch_sessionfor fork adoption on a live process. - Optional, for the next bug in this class: log Pi RPC command types (not payloads) around session bind; don’t start
get_commandsuntil the thread is bound.
Existing adapter test (“registers the thread from
get_stateand resumes viaswitch_session”) encodes the current cold-resume shape and will need to expect spawn-time--sessionplus no switch when the file already matches.No change needed on
main. Fix belongs ont3code/codex-turn-mapping.Workaround until then: resume the same session from a terminal with
pi --session <file>. Guardingpi.onin user extensions only hides their rows; T3’s hooks still no-op.- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.acceptedfeature request acceptedfeature request acceptedvia-triageFiled through npx t3 triageFiled through npx t3 triage
on Sep 18, 2026 - added a commit that references this issue
on Oct 2, 2026
What happened
Resumed a day-old Pi thread in T3 Code (work Mac, Orchestrator V2 branch build) and sent a message. The turn ran, but every Pi extension hook that touched its
ctxthrewThis extension ctx is stale after session replacement or reload…, rendered as red rows in the thread:cmux-session failed during before_agent_start(1×)index failed during before_agent_start(1×, a third-party extension)pi-t3-mcp-extension failed during before_provider_request(3× — once per provider call in the turn)The last one is T3's own injected extension (
piT3McpExtensionSource.ts). Itsbefore_provider_requesthandler reads only the hook-passedctx(ctx.model?.provider) and captures nothing across sessions, so the ctx handed to the handler was already invalidated — this is a dispatch-side problem, not an extension holding a stale reference. Every T3-owned hook on that turn silently no-op'd: the OpenRoutermax_tokenscap and, on other paths, thetool_callpermission gate.Resuming the same Pi session from a terminal (
pi --session <file>) and sending the same message: no errors.Diagnosis
T3 resumes a Pi thread by spawning
pi --mode rpc(no--session), which opens a default session and loads all extensions against it, then sendsswitch_sessionto the persisted file (PiAdapterV2.tsregisterThread,existing.nativeThreadRef.nativeId). In Pi 0.85.1 that switch is a full session replacement:teardownCurrent()→session.dispose()→extensionRunner.invalidate(...), thencreateRuntime()builds a fresh resource loader, re-runs every extension factory and installs a new runner. Any hook dispatched through the old runner after that point gets an invalidated ctx and throws on first access.Driving this exact sequence by hand against
pi --mode rpc(spawn →switch_session→prompt; also with the full global extension set; also after the default session had already run a turn) does not reproduce — the replaced runner is fresh and every hook is fine. So the plain switch is sound; the failing process had T3-specific history (kept alive across a day; whatever RPC sequence T3 issued to it between yesterday's last turn and today's resume) that left the dispatching runner invalidated while a turn still ran through it. I could not extract that sequence: the server trace does not record Pi RPC commands or the Pi child's stderr, andextension_errorrows are not persisted in the event store, so the evidence is the UI only.Two observations from the probes that may help whoever knows the pool lifecycle: (1)
session_startwithreason: "resume"fires twice perswitch_sessionin RPC mode —rpc-mode.tsrebinds once viaruntimeHost.setRebindSessioninsidefinishSessionReplacementand again incase "switch_session", and eachbindExtensionsemitssession_start; harmless in isolation, but it means every T3 resume double-runs every extension'ssession_start. (2) The whole class exists only because T3 replaces a session inside a live process. Spawning the resume aspi --mode rpc --session <file>(Pi's native resume;--sessionaccepts the file path T3 already stores asnativeThreadRef.nativeId) would load extensions once against the right session and removeswitch_sessionfrom the resume path entirely — no runner invalidation can then reach a turn. Every other CLI harness T3 drives is resumed by passing the session id on the command line; Pi supports the same.Steps to reproduce
Not deterministic yet. Observed twice (2026-09-10 and 2026-09-18) with the same shape: a Pi thread idle for ≥ 1 day, reopened in T3, first message sent. A raw-RPC replay of spawn →
switch_session→promptdoes not trigger it (see Diagnosis).Version
0.0.42 — Orchestrator V2 branch (
t3code/codex-turn-mapping), head934da5749, packaged desktop build.Environment
macOS 15 (arm64), Node 24, Pi
@earendil-works/pi-coding-agent0.85.1 on both machines (one reproduced, one is the probe machine).Evidence
UI rows (verbatim prefix; full text is Pi's standard stale-ctx message):
Pi source (0.85.1) for the replacement path:
dist/core/agent-session-runtime.jsswitchSession()→teardownCurrent()→dist/core/agent-session.jsdispose()line 595this._extensionRunner.invalidate(...);dist/modes/rpc/rpc-mode.jscase "switch_session"→rebindSession().Clean probe transcript (spawn → switch → prompt, probe extension logging the session file from the hook ctx):
Related issues
None found for
stale ctx/switch_session. #12285 (delegated-wake cap) and #11168 (mode:"wait") are different V2 contract problems from the same setup; not duplicates.Fix applied or workaround
User-side: every hand-built extension here now wraps
pi.onwith a guard that logs the stale throw once and no-ops afterwards (that is whycmux-sessionshows 1 row instead of one per call). That hides the symptom for our extensions only; T3's own extension and third-party ones still throw, and the guarded hooks still do nothing for the turn. Reliable recovery is to resume the session from a terminal instead of T3.Filed by
claude (opus-5) via a Pi thread inside T3 Code, following the
t3 triagereport structure.