Repository navigation
[Bug]: Nightly retries unreachable SSH environments much harder than Alpha (≈3× tunnel readiness failures, 2× local ports) #15616
Description
Activity
Note
Grok responding on behalf of Julius.
Triage
Thanks for the careful numbers, @smnbss. Having per-build counts from the trace files made this much easier to dig into.
What I found
I compared Alpha on 2 Oct (v0.0.45,
6c8fed35) with0.0.46-nightly.20261003.2623(fed41fa8). Between those two,packages/ssh/src/tunnel.tsonly changed imports, and the readiness cadence is the same: a probe every 100ms, up to 20s for a new forward (SSH_READY_TIMEOUT_MS), and a 2s check before an existing forward is dropped (ensureTunnelEntry).The relevant change in that range is #14897. When a connection keeps failing, reconnects now back off with jitter up to 5 minutes instead of stopping at 16 seconds. The first retry comes a little sooner (about 1–2s instead of 3s), and bringing the app to the foreground still resets the backoff, as it did on Alpha.
Here are your numbers converted to hourly rates, reading "until ~18:24" as 18.4 hours and the evening window as 6 hours:
Window Failed loopback GETs / hour Distinct local ports / hour Alpha, 2 Oct (24h) ~172 ~2.9 Alpha, 3 Oct until 18:24 ~221 ~5.4 Nightly, 3 Oct evening (~6h) ~1,220 ~24.5 Nightly, 4 Oct (24h) ~104 ~2.6 So the 3 Oct evening window is clearly hotter, but the full Nightly day on 4 Oct is quieter than Alpha. During the spike, about one new local port every 5 minutes per host (147 ports over 6h across 2 hosts) lines up with the new 5-minute cap rather than a tighter loop.
Also, those loopback GETs aren't SSH connection attempts.
launchOrReuseRemoteServerruns beforereserveLocalTunnelPort, so connect timeouts, MagicDNS failures, and the 15-minuteSSH command timed outnever show up ashttp.client GETto127.0.0.1. They come fromwaitForHttpReadyon a forward that already exists. A full 20s wait is about 200 probes, which is close to the ~195 per port you saw. The evening average works out to about 20 GETs per failedwaitForHttpReadyspan (7,317 / 369), which matches the 2s stale-tunnel check.This is related to #11287 and #4144 but isn't the same report.
To narrow down what made that evening hot, a short sanitized trace excerpt would help:
ssh/tunnel.launchOrReuseRemoteServerfailures that never reserve a port- each failed
shared.httpReadiness.waitForHttpReadyspan, with itstimeoutMsandattempts - whether an
application-activewakeup was resetting the backoff during the spike
The exact Alpha version string from that app would help too.
Likely fix area
- The reconnect backoff from fix(client-runtime): reconnects back off with jitter and keep healthy sockets #14897, especially how foreground and wakeup resets interact with it while a host stays unreachable.
- The stale-tunnel check in
ensureTunnelEntryandwaitForHttpReadyinpackages/ssh/src/tunnel.ts, if the excerpt shows forwards being recreated faster than the backoff allows.
A maintainer will decide on the fix direction.
Reacted by Simone Basso- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Oct 4, 2026
Before submitting
Area
Desktop app: SSH remote environments.
Steps to reproduce
T3 Code (Alpha)for a day, then switch to the Nightly build and run it for a day with the same environments.http.client GETspans against127.0.0.1:<port>in~/.t3/userdata/logs/*.trace.ndjson, grouped by day and by build (the stack trace includes the app bundle name).Expected behavior
When a remote environment is unreachable, the Nightly build should back off no faster than Alpha did, and should not start new tunnels/ports at a higher rate.
Actual behavior
The Nightly build makes far more failed readiness checks and uses far more distinct local tunnel ports than Alpha, against the same hosts:
http.client GETto127.0.0.10.0.46-nightly.20261003.26230.0.46-nightly.20261003.2623Most failures sit under the SSH environment path:
desktop.ipc.sshEnvironment.ensureEnvironment→ssh/tunnel.ensureEnvironment→ssh/tunnel.ensureTunnelEntry.create→ssh/tunnel.launchOrReuseRemoteServer/ssh/tunnel.startSshTunnelshared.httpReadiness.waitForHttpReady: 369 failuresdesktop:ensure-ssh-environmentIPC: 342 failures.desktop:bootstrap-ssh-bearer-session: 96 failures.The underlying SSH errors are real network failures, and both builds see them:
So the bug is not that SSH fails. The Nightly build seems to retry an unreachable environment much harder (more tunnels, more readiness polls) than Alpha. Each local port takes about 195 failed GETs before it is dropped.
No crashes: the traces have no
Die/Interruptexits, and macOS DiagnosticReports has no T3 crash reports.Impact
The app feels noticeably less stable on Nightly while a remote environment is unreachable. The extra churn also adds load on remote hosts, which is the same class of problem as #11287.
Version or commit
Nightly
0.0.46-nightly.20261003.2623, compared with T3 Code (Alpha) (version not recorded) on the same machine the day before.Environment
network-accessible, Tailscale Serve enabledLogs or stack traces
Representative readiness failure (Nightly):
Unrelated noise seen while triaging: 1,062
CommandResolutionErrorfailures fromshell.resolveCommandPath. These come fromexternalLauncher.resolveAvailableEditorsprobing for editors that are not installed. That looks expected, but it is very loud in the traces.Workaround
None found. Today's lower count may mean the 2026-10-03 spike was partly the switch to the new build. I can supply sanitized trace excerpts if useful.