Skip to content

[Bug]: Codex turn sees partial MCP tool catalog during server startup #13437

Description

@alvaromoran-aily

Area

apps/server — Codex provider

Version and environment

T3 Code desktop v0.0.42 on macOS; Codex with GPT-6 Luna. Several configured stdio MCP servers have different startup times.

Steps to reproduce

  1. Configure multiple Codex MCP servers, including one that initializes in about 10 seconds and others that take 20 seconds or longer.
  2. Start a fresh T3 Codex thread and immediately ask: “What MCPs do you have available?”
  3. Ask again in a follow-up turn if T3 starts a new provider session for it.

Expected behavior

The model can discover the configured MCP tools once their servers are ready, or receives a clear startup/incomplete-inventory state before answering. A tool inventory should not silently contain only the fastest servers.

Actual behavior

The model reported only the first two ready MCP servers and said it could not verify the rest. Other configured servers became ready before its answer, but their tools were absent from the catalog snapshot it used. In a separate Codex turn, a later tool-registry lookup found the missing tools and a call to one of them succeeded.

Relevant timing from a redacted provider trace

14:20:51.141  configured MCP servers: starting
14:20:51.219  turn/started
14:21:04.143  MCP server A: ready
14:21:04.447  MCP server B: ready
14:21:08.729  model calls list_mcp_resources
14:21:10–12   several more configured MCP servers: ready
14:21:20.877  model answers with only servers A and B
14:22:01–02   additional servers reach ready; two others fail

list_mcp_resources does not enumerate callable tools, which also makes the model's inventory unreliable. The timing issue is independent of those two failed servers: several successful servers reached ready before the answer but after the model's lookup.

In apps/server/src/provider/Layers/CodexSessionRuntime.ts, sendTurn calls config/mcpServer/reload and then turn/start. The reload response arrives before the mcpServer/startupStatus/updated notifications reach terminal states. Please account for that asynchronous startup before the model treats a tool catalog as complete. A bounded wait or a live catalog refresh could preserve responsiveness when a server is slow or fails.

Impact

Major degradation: Codex agents can skip working MCP tools or claim they are unavailable.

Workaround

A fresh tool-registry lookup after the server reaches ready can find the tools, but the model does not reliably perform one on its own.

Activity

  1. juliusmarminge commented on Sep 24, 2026

    @juliusmarminge
    Member

    Confirmed. This is still the case on main and in the reported v0.0.42 build. No existing issue covers it.

    CodexSessionRuntime.sendTurn treats config/mcpServer/reload as “the tool catalog is ready,” then immediately calls turn/start:

    https://github.com/pingdotgg/t3code/blob/b2b43bef7344/apps/server/src/provider/Layers/CodexSessionRuntime.ts#L2501-L2533

    That RPC does not mean startup has finished. Its response is an empty McpServerRefreshResponse, and the app-server only queues Op::RefreshMcpServers before returning. Each server then reports starting / ready / failed / cancelled on mcpServer/startupStatus/updated. Nothing in this runtime waits on those notifications, and it never calls mcpServerStatus/list. The session is already marked ready when the thread opens. mcpServer/startupStatus/updated is not mapped to a runtime event, so the thread gets no incomplete-inventory state.

    The reload runs on every turn where T3 attached mcp_servers.t3-code, not only the first one. It refreshes every configured server on the thread, including user stdio servers. Your trace matches that window: servers enter starting at 14:20:51.141 and turn/started arrives 78ms later, before any of them are ready.

    Codex then snapshots the model-facing tool binding with a 1s grace for optional servers (OPTIONAL_MCP_STARTUP_GRACE in codex-rs/codex-mcp/src/connection_manager/tool_catalog.rs). A server still starting after that is omitted (omitting pending optional MCP server). Required servers and codex_apps are waited on; ordinary configured servers are not. That lines up with a catalog that contains only the servers already ready at 14:21:04, while servers that reached ready at 14:21:10–12 stay out of that snapshot. A later turn can see them once startup has finished, which is why a follow-up registry lookup and tool call succeeded.

    list_mcp_resources is a separate Codex behavior: it lists resources, not callable tools (openai/codex#14242). It makes the model’s verbal inventory unreliable, but it does not explain tools that were already ready and still absent from the binding.

    A bounded wait before turn/start — until each expected server leaves starting, or a timeout elapses — would match what the Codex TUI does and would keep a hung or failed server from blocking the turn. Codex can emit a spurious cancelled before ready for a server that was only started once (openai/codex#36682), so the waiter has to key on server name and let a later ready win.

  2. added
    acceptedfeature request accepted
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Sep 24, 2026
  3. ScottN-PV commented on Oct 3, 2026

    @ScottN-PV
    Contributor

    This still happens on V2 main, and a Codex setting already closes the gap without a T3 change.

    CodexAdapterV2 sends turn/start a few milliseconds after thread/start returns. Codex waits mcp_optional_startup_grace_ms (default 1 s), then builds the first model request without any server that is still starting. It rebuilds the tool list for each later request in the turn, so a late server is missing only until it is ready.

    Workaround today: set mcp_optional_startup_grace_ms in the Codex home's config.toml. With 3000 on main, the first request waited 2.8 s and had every server. A value above the slowest server's start time should do the same for the 10–20 s servers in this report. I ran only 3000.

    Measured on Linux with Codex 0.159.3, agent-run, one run per case. Windows and macOS were not measured.

    • Warm npx/uvx servers (filesystem, memory, sequential-thinking, Playwright, time, fetch) were ready 0.4–0.7 s after turn/start. None was dropped.
    • With empty package caches they were ready after 1.4–2.7 s, and all six were left out of the first request.
    • In the repo's recorded transcripts, xcodebuildmcp takes 2.5–5.0 s to start (median 4.0 s, 29 recordings). In the 23 recordings where it starts with the thread, it is ready 3.8–5.5 s after the turn starts (median 4.6 s).
    • With grace 3000 and one server that never answers, the first request waited 3.0 s and left out only that server. The second turn did not wait.

    Do you want T3 to set a default, or leave this to the Codex setting?

    1. Docs only. Describe the setting in docs/user/providers-codex.md. No default changes.
    2. A finite grace in T3's thread config, for example 5 s. It covers cold npx starts and 18 of those 23 recorded xcodebuildmcp starts, but not the 10–20 s servers in this report. A hung server would add about 4 s to the first turn of each Codex session (not run at 5 s).
    3. 0 in T3's thread config. Codex then waits for each server up to its startup_timeout_sec, 30 s per startup phase by default. A server that never answered held the first turn about 30 s.

    Options 2 and 3 would override the user's own config.toml value. That comes from the Codex source and was not run.

    I can send whichever you choose.

    Written by Claude Fable 5.1 on behalf of @ScottN-PV.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions