Skip to content

[Bug]: t3connect mobile session drops every ~4.5s on Windows 0.0.31; server.getConfig still blocked by editor discovery (follow-up to #3610) #5137

Description

@varun-101

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. Windows 11, T3 Code desktop 0.0.31, with a large PATH (46 entries) and the default 14-entry PATHEXT, and most of the JetBrains IDEs in EDITORS not installed.
  2. Set serverExposureMode: "network-accessible" and connect a phone to the desktop environment over t3connect (relay).
  3. The websocket upgrade authenticates successfully, then the session drops after ~4-5 seconds.
  4. The phone reconnects and the cycle repeats indefinitely.

Expected behavior

The mobile session connects and stays connected. Optional editor discovery should not be able to hold the first server.getConfig for multiple seconds on the connection-critical path.

Actual behavior

This is a follow-up to #3610 (closed as completed via #4291). The Effect.timeoutOption bound added in #4291 is present in 0.0.31 and does fire, but a 5 second bound is still far longer than the mobile session survives, so the reconnect loop persists.

In every websocket I traced, the connection lives exactly as long as server.getConfig takes, then SessionStore.markDisconnected fires immediately:

/ws lifetime ws.rpc.server.getConfig externalLauncher.buildAvailableEditors
4547 ms 4325 ms 4325.08 ms — Success
4489 ms 4381 ms 4381.02 ms — Success
5165 ms 5001 ms 5000.31 ms — Interrupted (hit EDITOR_DISCOVERY_TIMEOUT)

Auth itself is healthy and fast: authenticateWebSocketUpgrade 0.96 ms, markConnected 1.83 ms. The entire connection lifetime is editor discovery.

I want to be precise about what I did and did not establish: I confirmed the correlation and the ordering (markDisconnected immediately follows getConfig completing, in all three connections). I did not isolate which side actually closes the socket.

Why discovery takes 4-5 seconds

Two compounding things in apps/server/dist/bin.mjs (0.0.31 bundle):

1. Candidate list is doubled for no benefit on Windows. resolveCommandCandidates pushes both the upper- and lower-case form of every PATHEXT entry, on a case-insensitive filesystem:

// bin.mjs:15741-15746
const candidates = [];
for (const candidateExtension of windowsPathExtensions) {
	candidates.push(`${command}${candidateExtension}`);
	candidates.push(`${command}${candidateExtension.toLowerCase()}`);
}
return Array.from(new Set(candidates));

With the default 14-entry PATHEXT that is 28 candidates instead of 14. resolveCommandPathForPlatform then loops for (const pathEntry of pathEntries) for (const candidate of commandCandidates), so on this machine each lookup is 46 dirs x 28 candidates = 1,288 fileSystem.stat calls. Exactly half of those are redundant.

2. No caching, and misses cost the most. buildAvailableEditors walks all 21 entries of EDITORS on every call. Installed editors early-exit, but the ~15 uninstalled JetBrains IDEs each perform a complete PATH sweep. Nothing is memoized, so this re-runs on every getConfig, i.e. on every connection attempt.

Measured in my traces: 28,511 shell.isExecutableFile spans per 10 MB trace file.

Secondary effect: trace log flooding

Because each failed connection triggers a fresh full sweep and the client immediately retries, the loop is self-sustaining and generates enormous trace volume. server.trace.ndjson rotated 10 files x ~10 MB in about 5 minutes (23:19 - 23:24) while the phone was retrying. shell.isExecutableFile spans are ~98% of that volume.

Possibly related

#4697 ("No installed editors found", 0.0.30-nightly, Windows 11) looks like the visible symptom of the same 5 s bound expiring and Option.getOrElse(() => []) returning an empty editor list.

Impact

Blocks work completely

Version or commit

0.0.31 (desktop, Alpha)

Environment

Windows 11 Home Single Language 26200, T3 Code desktop 0.0.31, serverExposureMode: network-accessible, connecting via t3connect relay from Android. PATH: 46 entries. PATHEXT: 14 entries (.COM;.EXE;.BAT;.CMD;.VBS;.VBE;.JS;.JSE;.WSF;.WSH;.MSC;.PY;.PYW;.CPL).

Logs or stack traces

# Secrets redacted: wsTicket JWT, relay subdomain, session id, client IP, username.
# One full websocket lifecycle, trace 8795e9b1f8591bd5...  (server.trace.ndjson)

name: EnvironmentAuth.authenticateWebSocketUpgrade   durationMs: 0.9586     exit: Success
name: SessionStore.markConnected                     durationMs: 1.8313     exit: Success
name: ws.rpc.server.getConfig                        durationMs: 5001.3239  exit: Success
name: SessionStore.markDisconnected                  durationMs: 0.8873     exit: Success
name: http.server GET  route=/ws                     durationMs: 5165.1161  status: 204

# Inside that getConfig:
name: externalLauncher.resolveAvailableEditors       durationMs: 5000.7373  exit: Interrupted
name: externalLauncher.buildAvailableEditors         durationMs: 5000.3134  exit: Interrupted
  -> EDITOR_DISCOVERY_TIMEOUT (5s, bin.mjs:43599) fired; availableEditors returned []

# The two preceding connections completed under the bound and still dropped:
name: ws.rpc.server.getConfig                        durationMs: 4325.6193  exit: Success
name: externalLauncher.buildAvailableEditors         durationMs: 4325.0778  exit: Success
name: http.server GET  route=/ws                     durationMs: 4547.0357

name: ws.rpc.server.getConfig                        durationMs: 4381.6312  exit: Success
name: externalLauncher.buildAvailableEditors         durationMs: 4381.0179  exit: Success
name: http.server GET  route=/ws                     durationMs: 4489.1237

# Volume: 28,511 shell.isExecutableFile spans in a single 10 MB trace file.
# 10 such files rotated in ~5 minutes during the reconnect loop.

# Repeated failing lookups for uninstalled editors:
name: shell.resolveCommandPathForPlatform  exit: Failure  cause: CommandResolutionError
name: shell.resolveCommandPath             exit: Failure  cause: CommandResolutionError

Screenshots, recordings, or supporting files

No response

Workaround

Reducing the per-lookup probe count on the client machine is enough to get under the threshold. On my machine, trimming PATHEXT from 14 entries to .COM;.EXE;.BAT;.CMD cuts candidates per lookup from 28 to 8, and removing 6 dead PATH entries (stale JetBrains dirs on a removed path, an unexpanded %MAVEN_HOME%\bin, and two entries that point at .exe files rather than directories) takes it from 1,288 stats per lookup to roughly 320.

This is a user-side mitigation for a server-side cost problem, and it will not help anyone who genuinely needs a large PATH.

Suggested fixes, roughly in increasing order of effort:

  1. Drop the lower-case duplicates in resolveCommandCandidates on win32. Windows path resolution is case-insensitive, so this halves the syscall count for a one-line change.
  2. Memoize editor discovery for the lifetime of the server process, or invalidate on PATH change. It currently re-runs on every getConfig.
  3. Lower EDITOR_DISCOVERY_TIMEOUT. 5 s is longer than a mobile session tolerates. Something in the 500-750 ms range would fail fast into the existing empty-list fallback.
  4. Move discovery off the getConfig critical path entirely and push results to the client when they arrive, which is the direction discussed in [Bug]: Windows 0.0.28/nightly local backend appears disconnected; server.getConfig interrupted, 0.0.27 works #3610. That would make the timeout value irrelevant rather than just smaller.

Activity

  1. CDVolvik commented on Aug 11, 2026

    @CDVolvik
    Contributor

    Re-checked both of your root causes against today's main. One of them has already landed.

    Cause 2, no caching, is fixed. packages/shared/src/shell.ts now carries a real CommandResolutionCache — a Context.Reference<Map> with a 30s TTL on the monotonic clock and a 512-entry cap, injectable so tests get an isolated instance (shell.ts:505-540). Worth flagging because you read the shipped dist/bin.mjs, and the bundle lags source.

    Cause 1, doubled candidates, is still live, at shell.ts:488-491: every PATHEXT entry is pushed twice, once as-is and once lower-cased. Your arithmetic holds.

    It is not a two-line deletion, though. parseWindowsPathExtensions (shell.ts:461) upper-cases every extension first, so the original PATHEXT casing is already gone by the time candidates are built. The two probes are foo.CMD and foo.cmd, and dropping either is a behaviour change rather than a pure saving:

    • drop the lower-case probe, and a real foo.cmd goes invisible in any directory with NTFS per-directory case sensitivity enabled, including every \wsl.localhost\... path
    • drop the upper-case probe, and an actual FOO.CMD in such a directory goes invisible

    Case-insensitive is the norm on win32, so either choice is almost always safe, and almost always is how something ships that reproduces for one person and nobody else. Real files are lower-case, so keeping the lower-case probe looks like the defensible pick — but that is a judgement about which filesystems you intend to support, not one I should make in a PR.

    One caveat on the symptom, which is the reason I would rather ask than just send a speedup: EDITOR_DISCOVERY_TIMEOUT is 5s (apps/server/src/ws.ts:129), still above the ~4.5s session survival you measured. Halving the probe count may not on its own fix the disconnect.

    Happy to implement whichever probe you want kept, with a test pinning the case-sensitive directory case.

  2. t3dotgg commented on Aug 27, 2026

    @t3dotgg
    Member

    Thanks for the detailed report. We believe this is fixed by PR #5561 and PR #5572.

    The report measures repeat editor scans of 4.3 to 5.0 seconds followed by disconnects. The landed reconnect repair makes the RPC heartbeat tolerate two missed 5-second windows, caches command lookup for 30 seconds, and caches successful editor discovery for 60 seconds. The follow-up repair prevents an interrupted scan from poisoning that cache. These changes remove the measured reconnect cycle even though initial discovery can still wait up to five seconds. The later issue comment reviews source but does not report a failed run with the new client.

    I'm closing this as fixed as part of an automated pass on all open issues. If this still happens in a current build that includes PR #5561 and PR #5572, please reply with the T3 Code version and fresh logs or reproduction details, and we can reopen it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions