Skip to content

[Bug]: Windows: Codex provider probe exceeds the 10 s AUTH_PROBE_TIMEOUT on ~1 in 4 runs and flips the card to error until the next refresh #7513

Description

@krutftw

What happened

On Windows the Codex provider card intermittently flips to error — toast "Codex provider status: Timed out while checking Codex app-server provider status." — even though Codex is installed, logged in (ChatGPT Pro) and works. It stays in that state until the next health refresh (5 min by default), and it happens most reliably right after app start or a renderer reload, when all providers probe at once.

Diagnosis

checkCodexProviderStatus (apps/server, //#region src/provider/Layers/CodexProvider.ts in the shipped bundle) runs probeCodexAppServerProvider — spawn codex app-server (via the npm .cmd shim on Windows → node → native exe), initialize, account/read, then skills/list + the full model listing — inside a single timeoutOption(AUTH_PROBE_TIMEOUT_MS) with AUTH_PROBE_TIMEOUT_MS = 1e4 (10 s, src/provider/providerSnapshot.ts). On timeout it returns status: "error", auth: unknown, models = custom-only, discarding the previous good snapshot.

Measured on this machine today (server.trace.ndjson, probeCodexAppServerProvider / checkCodexProviderStatus):

time (UTC) probe check outcome
08:05:12 4,990 ms 5,599 ms ok
08:05:23 9,794 ms 11,324 ms timed out (boot)
08:10:12 2,512 ms 3,144 ms ok
08:15:15 2,451 ms 2,883 ms ok
08:20:18 1,696 ms 2,135 ms ok
08:30:20 2,481 ms 2,917 ms ok
08:45:34 4,109 ms 4,865 ms ok
08:53:16 10,068 ms (Interrupted) 12,072 ms timed out (renderer reload)

2 of 8 probes exceed the 10 s bound; the others take 1.7–5 s, so the bound is within normal variance for this platform under load (5 providers probing concurrently + the known 5 s editor-discovery block, #4697).

Suggested fix (small): either raise AUTH_PROBE_TIMEOUT_MS on Windows (or overall), or on timeout keep the last known-good snapshot and mark it stale (status "warning", "Codex probe timed out; showing last known status") instead of flipping to error with an empty model list.

Steps to reproduce

  1. Windows desktop app, Codex installed via npm (@openai/codex 0.148.0, ChatGPT login).
  2. Start the app (or reload the renderer) so all provider probes run concurrently.
  3. Watch caches/codex.json / the Codex card: roughly 1 in 4 probes lands at status: "error", message "Timed out while checking Codex app-server provider status.", until the next 5-minute refresh.

Version

0.0.34-nightly.20260819.1132 (commit 36f4314).

Environment

Windows 11 Pro 10.0.26100, Node 24.11.1, @openai/codex 0.148.0 (npm, via %APPDATA%\npm\codex.cmd), ChatGPT Pro login.

Evidence

# caches/codex.json after the 08:53:16 probe
{"status":"error","installed":true,"version":null,"auth":{"status":"unknown"},
 "message":"Timed out while checking Codex app-server provider status.","checkedAt":"2026-08-19T08:53:16.976Z"}

# bundle
const AUTH_PROBE_TIMEOUT_MS = 1e4;
... probe({...}).pipe(scoped$1, timeoutOption(millis(AUTH_PROBE_TIMEOUT_MS)), result$1)
... if (isNone(probeResult.success)) return buildServerProvider({ ... status: "error", auth: { status: "unknown" }, message: "Timed out while checking Codex app-server provider status." })

Related issues

#4697 / #5137 (editor discovery blocking server.getConfig for 5 s on Windows) is the likely source of the concurrent load, but this is about the provider probe's own bound and error handling; no existing issue mentions the Codex probe timeout.

Fix applied or workaround

None locally (the timeout is a constant). It self-heals at the next 5-minute refresh or on a settings change.

Filed by

Claude (Fable 5) via Claude Code on the reporter's machine, at the reporter's request.

Activity

  1. 123Kazukiouchi commented on Oct 4, 2026

    @123Kazukiouchi

    I can reproduce this on Windows 11 with T3 Code 0.0.46-nightly.20261003.2638 and Codex CLI 0.160.0. Codex works normally from the CLI, and Pi works inside T3. A manual Codex provider refresh after the Pi task had finished still timed out.

    In three failed T3 checks, probeCodexAppServerProvider was interrupted at 10,058 / 10,082 / 10,092 ms. Initialization completed in 2,011–3,297 ms, and model listing completed in 8–365 ms.

    In this release, the rate-limit request has its own 3,000 ms timeout. Based on those trace spans and the probe's request order, the skills/list path appears to have still been outstanding when the overall 10-second provider-probe deadline was reached. This part is an inference because the T3 trace does not record a separate skills/list completion timestamp.

    I also ran three independent Codex app-server probes using T3's initialize parameters and concurrent requests:

    Stage Run 1 Run 2 Run 3
    Start + initialize 1,525 ms 1,759 ms 1,734 ms
    account/read 451 ms 350 ms 281 ms
    skills/list 4,924 ms 7,757 ms 5,593 ms
    model/list 4 ms 18 ms 3 ms
    account/rateLimits/read 736 ms 457 ms 418 ms
    Total 6,917 ms 9,884 ms 7,622 ms

    The independent probe used the native codex.exe, so its totals are not identical to T3's npm-shim/provider path. Even so, one successful run finished only 116 ms below T3's 10-second overall budget.

    These measurements point to normal variation in the skills/list path pushing an otherwise healthy full provider probe over the fixed 10-second budget. They do not indicate a Codex CLI failure.

    This seems to support either increasing the Windows Codex provider-probe budget or preserving the last known-good provider state when a probe times out, rather than marking Codex unavailable.

  2. added 2 commits that reference this issue on Oct 8, 2026
    ccb5cc4
    9466c36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions