Repository navigation
[Bug]: SSH environment stays reconnecting after desktop update reuses an older remote server #16984
Description
Activity
Note
Grok responding on behalf of Julius.
Thanks for the detailed report. Your logs point to a specific mechanism in the code. When the desktop and the reused remote server are on different builds, the pairing token is minted by one build and redeemed by the other. Reading the code, an Oct 4 server can't read tokens minted by an Oct 7 build.
What the reconnect does (main @
300f7f9)- The launch script runs the runner once, which installs the desktop's pinned archive version. That's why the Oct 7 runtime showed up on the remote (tunnel.ts#L568-L573).
- If
~/.t3/userdata/server-runtime.jsonnames a live server whose port answers the readiness probe, that server is adopted asexternaland reused. Nothing compares its version to the desktop's (tunnel.ts#L630-L672). Only themanagedbranch restarts a server when the runner script changes (#L673-L686). - The pairing token is then minted by running
auth pairing create --base-dir ~/.t3through the new runner. It is written to the shared database, not issued by the running server (tunnel.ts#L723-L734, #L1689-L1697). It gets the defaultAuthStandardClientScopes(cli/auth.ts#L85-L104).
Why the Oct 4 server returns 500
- On Oct 7 (03:30 UTC), the permission split (feat(auth): separate environment administration permissions #9786 to feat(auth): allow passive terminal observation #9791, fix(auth): keep old clients connected across scope changes #10298) added new scopes:
settings:write,providers:manage,environment:maintain,preview:operate,diagnostics:read,terminal:read,source-control:write,filesystem:readandfilesystem:write. All of them are now in the standard client scopes (contracts/auth.ts#L205-L219). - The Oct 4 build (
v0.0.46-nightly.20261004.2652=4ee6bfd) only knows 8 scope literals (auth.ts#L89-L98). Its pairing-link row decodesscopesstrictly withSchema.fromJsonString(AuthEnvironmentScopes)(AuthPairingLinks.ts#L18-L30). - So
consumeAvailablefinds the row, but decoding itsRETURNINGrow fails with aPersistenceDecodeError(#L261-L283). That error is wrapped asBootstrapCredentialConsumeAvailableError(PairingGrantStore.ts#L512-L519), then becomesServerAuthBootstrapCredentialValidationError(EnvironmentAuth.ts#L531-L541) and finally a 500access_token_issuance_failed(http.ts#L381-L383). That is exactly the error chain in your log. - fix(auth): keep old clients connected across scope changes #10298 handled older clients talking to newer servers. It didn't cover a newer CLI writing grants that an older server has to decode.
- Each reconnect mints a fresh token the same way, so the loop never recovers. Restarting the server on the Oct 7 build fixes it because that build knows every scope.
This was confirmed by reading the code. I haven't reproduced it end to end.
Not #16797. The boolean bind (
${requestedScopes === undefined}) arrived with #10298 and isn't in the Oct 4 query (AuthPairingLinks.ts @ 4ee6bfd). The Oct 7 server that worked for you does have it.Why the UI only says "reconnecting": every SSH preparation failure that isn't a cancellation, including this 500, is mapped to
ConnectionTransientError("remote-unavailable")(platform.ts#L170-L182). That error goes into backoff, and backoff shows asreconnecting(presentation.ts#L45-L50). The chat banner title is just " is reconnecting" (ChatView.tsx#L3118).Fix direction
- Detect a version mismatch on reuse. Compare the reused server's version with the runner's archive version. Then either restart it on the matching build or block with a clear "remote server is X, desktop expects Y" message. Avoid killing busy servers ([Bug]: SSH reconnect kills a busy remote server, cancelling all running agents (2s reuse probe) #16477).
- Alternatively, mint the token through the running server, or request only scopes it understands, so the issuer and consumer always agree.
- Forward compatibility: drop unknown scope strings when decoding stored grants instead of failing. That prevents the next scope addition from causing a 500 on older servers, though it can't fix builds already shipped.
- Treat a 5xx from
bootstrap-bearer-sessionas a blocking error that shows the reason, not a silent transient retry.
Workaround: restart the remote T3 server on the build that matches the desktop. That's what you did, keeping the same
--base-dirand port.Related: #16797 / #16730 (a different consume-query failure), #16477, #5749, #13341 and #13521 (SSH reuse/restart behavior; more reuse means more chances of version skew), #13717 and #11287 (other ways to get stuck reconnecting).
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Oct 7, 2026
Before submitting
Area
apps/desktop
Steps to reproduce
This is the sequence observed on my setup; I haven't reproduced it from a clean installation.
0.0.46-nightly.20261004.2652.0.0.46-nightly.20261007.2787.Expected behavior
The existing environment reconnects. If the remote server needs an update, T3 handles that safely or explains what action is required.
Actual behavior
The environment stayed on “ is reconnecting” (host name redacted). SSH and the forwarded HTTP server were reachable, but the desktop's pairing-token exchange returned HTTP 500.
The October 7 runtime was already installed remotely, while the process serving the SSH environment was still running the October 4 build. The UI didn't explain the authentication failure.
Impact
Blocks work completely in the affected remote environment.
Version or commit
Desktop:
0.0.46-nightly.20261007.2787Remote server still running:
0.0.46-nightly.20261004.2652Environment
Linux desktop and Linux SSH server, SSH over Tailscale.
Logs or stack traces
Workaround
After backing up the database and checking for active provider sessions, I restarted only the remote T3 server using
0.0.46-nightly.20261007.2787, keeping the same base directory and port.The desktop then reconnected successfully. Live RPC requests succeeded, the reconnecting message disappeared, and the existing chat history remained visible.
This points to reuse of the older server as the trigger, though I haven't isolated the underlying token-consumption failure.
Related reports