Skip to content

[Bug]: Nightly retries unreachable SSH environments much harder than Alpha (≈3× tunnel readiness failures, 2× local ports) #15616

Description

@smnbss

Before submitting

Area

Desktop app: SSH remote environments.

Steps to reproduce

  1. On macOS, add two SSH remote environments. Both hosts are reached over Tailscale (MagicDNS names), and each is sometimes unreachable (asleep or on a flaky link).
  2. Run T3 Code (Alpha) for a day, then switch to the Nightly build and run it for a day with the same environments.
  3. Count failed http.client GET spans against 127.0.0.1:<port> in ~/.t3/userdata/logs/*.trace.ndjson, grouped by day and by build (the stack trace includes the app bundle name).

Expected behavior

When a remote environment is unreachable, the Nightly build should back off no faster than Alpha did, and should not start new tunnels/ports at a higher rate.

Actual behavior

The Nightly build makes far more failed readiness checks and uses far more distinct local tunnel ports than Alpha, against the same hosts:

Build Window Failed http.client GET to 127.0.0.1 Distinct local ports
Alpha 2026-10-02, full day 4,117 69
Alpha 2026-10-03, until ~18:24 4,076 100
Nightly 0.0.46-nightly.20261003.2623 2026-10-03, ~18:24 → midnight (~6 h) 7,317 147
Nightly 0.0.46-nightly.20261003.2623 2026-10-04, full day 2,502 63

Most failures sit under the SSH environment path:

  • desktop.ipc.sshEnvironment.ensureEnvironment → ssh/tunnel.ensureEnvironment → ssh/tunnel.ensureTunnelEntry.create → ssh/tunnel.launchOrReuseRemoteServer / ssh/tunnel.startSshTunnel
  • shared.httpReadiness.waitForHttpReady: 369 failures
  • desktop:ensure-ssh-environment IPC: 342 failures. desktop:bootstrap-ssh-bearer-session: 96 failures.

The underlying SSH errors are real network failures, and both builds see them:

SshCommandError: ssh: connect to host <host> port 22: Operation timed out
SshCommandError: ssh: Could not resolve hostname <host>.<tailnet>.ts.net: nodename nor servname provided, or not known
SshCommandError: SSH command timed out after 900000ms.

So the bug is not that SSH fails. The Nightly build seems to retry an unreachable environment much harder (more tunnels, more readiness polls) than Alpha. Each local port takes about 195 failed GETs before it is dropped.

No crashes: the traces have no Die/Interrupt exits, and macOS DiagnosticReports has no T3 crash reports.

Impact

The app feels noticeably less stable on Nightly while a remote environment is unreachable. The extra churn also adds load on remote hosts, which is the same class of problem as #11287.

Version or commit

Nightly 0.0.46-nightly.20261003.2623, compared with T3 Code (Alpha) (version not recorded) on the same machine the day before.

Environment

  • macOS 27 (Darwin 27.0.0), Apple Silicon
  • Server exposure mode network-accessible, Tailscale Serve enabled
  • Two SSH remote environments (one macOS host, one Linux host), reached via Tailscale MagicDNS / Tailscale IPv6

Logs or stack traces

Representative readiness failure (Nightly):

HttpClientError: Transport error (GET http://127.0.0.1:50705/)
    at catch (/Applications/T3 Code (Nightly).app/Contents/Resources/app.asar/...)

Unrelated noise seen while triaging: 1,062 CommandResolutionError failures from shell.resolveCommandPath. These come from externalLauncher.resolveAvailableEditors probing for editors that are not installed. That looks expected, but it is very loud in the traces.

Workaround

None found. Today's lower count may mean the 2026-10-03 spike was partly the switch to the new build. I can supply sanitized trace excerpts if useful.

Activity

  1. juliusmarminge commented on Oct 4, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks for the careful numbers, @smnbss. Having per-build counts from the trace files made this much easier to dig into.

    What I found

    I compared Alpha on 2 Oct (v0.0.45, 6c8fed35) with 0.0.46-nightly.20261003.2623 (fed41fa8). Between those two, packages/ssh/src/tunnel.ts only changed imports, and the readiness cadence is the same: a probe every 100ms, up to 20s for a new forward (SSH_READY_TIMEOUT_MS), and a 2s check before an existing forward is dropped (ensureTunnelEntry).

    The relevant change in that range is #14897. When a connection keeps failing, reconnects now back off with jitter up to 5 minutes instead of stopping at 16 seconds. The first retry comes a little sooner (about 1–2s instead of 3s), and bringing the app to the foreground still resets the backoff, as it did on Alpha.

    Here are your numbers converted to hourly rates, reading "until ~18:24" as 18.4 hours and the evening window as 6 hours:

    Window Failed loopback GETs / hour Distinct local ports / hour
    Alpha, 2 Oct (24h) ~172 ~2.9
    Alpha, 3 Oct until 18:24 ~221 ~5.4
    Nightly, 3 Oct evening (~6h) ~1,220 ~24.5
    Nightly, 4 Oct (24h) ~104 ~2.6

    So the 3 Oct evening window is clearly hotter, but the full Nightly day on 4 Oct is quieter than Alpha. During the spike, about one new local port every 5 minutes per host (147 ports over 6h across 2 hosts) lines up with the new 5-minute cap rather than a tighter loop.

    Also, those loopback GETs aren't SSH connection attempts. launchOrReuseRemoteServer runs before reserveLocalTunnelPort, so connect timeouts, MagicDNS failures, and the 15-minute SSH command timed out never show up as http.client GET to 127.0.0.1. They come from waitForHttpReady on a forward that already exists. A full 20s wait is about 200 probes, which is close to the ~195 per port you saw. The evening average works out to about 20 GETs per failed waitForHttpReady span (7,317 / 369), which matches the 2s stale-tunnel check.

    This is related to #11287 and #4144 but isn't the same report.

    To narrow down what made that evening hot, a short sanitized trace excerpt would help:

    • ssh/tunnel.launchOrReuseRemoteServer failures that never reserve a port
    • each failed shared.httpReadiness.waitForHttpReady span, with its timeoutMs and attempts
    • whether an application-active wakeup was resetting the backoff during the spike

    The exact Alpha version string from that app would help too.

    Likely fix area

    A maintainer will decide on the fix direction.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions