Skip to content

[Bug]: One transient probe timeout leaves Claude stuck on "Could not verify Claude authentication status" for hours (cause swallowed, no demand-less recheck) #13635

Description

@oliwer-cpu

Area

apps/server (provider status / Claude capability probe)

Summary

On a headless Linux server (t3 serve, nightly 0.0.43-nightly.20260922.2110), one transient host stall made every provider check fail at once. Those failed snapshots then stayed in place for 22+ hours. The Claude card and toast kept saying "Could not verify Claude authentication status from initialization result." even though Claude was logged in and working the whole time.

Evidence

~/.t3/caches/*.json at 2026-09-25T11:02Z:

claudeAgent.json  checkedAt 2026-09-24T12:59:28Z  warning  Could not verify Claude authentication status from initialization result.
codex.json        checkedAt 2026-09-24T13:00:40Z  error    Timed out while checking Codex app-server provider status.
cursor.json       checkedAt 2026-09-24T13:00:07Z  error    Cursor Agent CLI is installed but timed out while running `agent about`.
  • claude auth status reports loggedIn: true (claude.ai subscription). Claude threads kept running normally in T3.
  • I reproduced the probe outside T3 with the same options (buildClaudeCapabilitiesProbeQueryOptions equivalents: empty settingSources, disableAllHooks, no MCP, cwd = server cwd), the bundled SDK version (@anthropic-ai/claude-agent-sdk@0.3.276) and the service's minimal env. initializationResult() returned the account in 1.3–2.0 s on 6 of 6 runs.
  • During a 5.5-minute watch, the server did not spawn a single capability probe, although providerHealthRefreshInterval is 1 min (performance profile).

Why it sticks (reading the bundled server)

  1. checkClaudeProviderStatus calls resolveCapabilities(...).pipe(orElseSucceed(() => undefined)), and probeClaudeCapabilities maps any failure to undefined. The real cause (timeout, spawn error, SDK error) is discarded and never logged, so the card can't say what went wrong.
  2. The periodic refresh only runs when backgroundPolicy.shouldRunScopeWork({ type: "provider-status" }) is true, meaning a client lease with provider-status demand exists and the host isn't constrained. If no client asks for provider status, a snapshot that failed transiently is never re-checked. It is also persisted to ~/.t3/caches/<provider>.json, so every reconnecting client sees it as current.

Expected

  • A failed or timed-out capability probe is retried on a short backoff whether or not a client has demand. At minimum, it is re-checked the next time a client connects.
  • The card/toast includes the underlying cause (e.g. "capability probe timed out after 25 s"). An unverified auth status after a probe failure shouldn't read like an auth problem.
  • A stale snapshot shows its checkedAt age in the UI.

Related

#7111 / #7175: same failure point, but for the slash-command list. That fix preserves commands, while the auth status still sticks. #7513: a similar probe timeout on the Codex provider.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions