Skip to content

fix(connections): avoid provider refresh storms on reconnect - #11456

Closed
Bil0000 wants to merge 2 commits into
pingdotgg:mainfrom
Bil0000:fix/reconnect-provider-refresh
Closed

Bil0000 wants to merge 2 commits into
pingdotgg:mainfrom
Bil0000:fix/reconnect-provider-refresh

Conversation

@Bil0000

@Bil0000 Bil0000 commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor

Each server-config subscription started a full provider refresh, so reconnects and extra tabs repeated provider probes. The server change removes that redundant refresh and keeps cached snapshots and provider-owned live updates.

Julius split this exact server patch into #11811, which merged at 7931227. This PR is now superseded by that merged fix.

Following the review from Julius, commit 0237d71 removes the proposed 45-second setup timeout and restores the existing 15-second timeout and its regression tests. The removed slow-setup tests used artificial delays; they did not establish a real connection that required a longer deadline. Connection readiness follows WebSocket connection, not shell sync.

Verified: 58 connection supervisor/RPC session tests and 8 server-config subscription tests pass. Client-runtime typecheck, scoped lint, formatting, and Ponytail review pass. A merge preview against main at 7931227 produces the exact same tree as main, with no conflicts or new changes.

There is no remaining change to land from this PR. No new PR or stack was created.

Model: GPT-6. Harness: Codex.

@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XS 0-9 changed lines (additions + deletions). labels Sep 12, 2026
@macroscopeapp

macroscopeapp Bot commented Sep 12, 2026 •

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Approved at 0237d71

Macroscope's review found this PR approvable — The production change removes redundant provider probes during WebSocket reconnects while preserving cached snapshots and independent background updates. A focused regression test covers repeated connections and verifies that no refresh is started.

Notes:

  • No code objects were reviewed. Approvability was decided on eligibility alone.

You can add or adjust custom eligibility rules. Learn more.

@coderabbitai

coderabbitai Bot commented Sep 12, 2026 •

Copy link
Copy Markdown

Review Change StackReview Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 7f254afa-07a9-4ff9-bdc3-661fb61cae2f

📥 Commits

Reviewing files that changed from the base of the PR and between c542b78 and fbdda7f.

📒 Files selected for processing (4)
  • apps/server/src/server.test.ts
  • apps/server/src/ws.ts
  • packages/client-runtime/src/connection/supervisor.test.ts
  • packages/client-runtime/src/connection/supervisor.ts
💤 Files with no reviewable changes (1)
  • apps/server/src/ws.ts

Included review availability: Your plan provides up to 10 included reviews per hour; 8 remain after this review.


📝 Walkthrough

Walkthrough

The server subscription no longer triggers provider refreshes. The connection supervisor now allows 45 seconds for establishment, with tests covering slow setup, reconnects, and timeout behavior.

Changes

Server configuration subscription

Layer / File(s) Summary
Snapshot-based subscription flow
apps/server/src/ws.ts, apps/server/src/server.test.ts
The subscription handler no longer starts providerRegistry.refresh(). Tests verify that reconnecting subscriptions receive snapshots without provider refresh calls.

Connection establishment timeout

Layer / File(s) Summary
Extended establishment window and coverage
packages/client-runtime/src/connection/supervisor.ts, packages/client-runtime/src/connection/supervisor.test.ts
The establishment timeout increases from 15 to 45 seconds. Tests cover slow setup, reconnects, stalled relay connections, timeout release, and platform wakeups.

Priority: ⬇️ Low

Estimated code review effort: 3 (Moderate) | ~25 minutes

Change: Bug fix

Suggested reviewers: t3dotgg

Merge Risk: ⚪ Minimal · up to 0237d

The subscription path now serves cached snapshots without triggering provider refreshes, and the longer establishment timeout is covered for slow setup, reconnect, timeout, relay, and wakeup behavior.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 3…
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly identifies the primary change: preventing redundant provider refreshes during reconnects.
Description check ✅ Passed The description explains the provider refresh change, the timeout decision, test results, and superseded status. It does not use the template headings or checklist, but it contains the required inform…
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@juliusmarminge

Copy link
Copy Markdown
Member

Audited fbdda7f. The server half is right and I split it out so it can land on its own: #11811 carries the ws.ts and server.test.ts hunks with you as commit author. Managed providers already run their own startup probe and periodic refresh, and the subscription stream seeds current providers before live changes, so the per-subscribe refresh bought nothing.

That leaves this PR with the CONNECTION_ESTABLISHMENT_TIMEOUT bump from 15s to 45s, and I don't think it should land as is. That timeout covers resolver.prepare → socket open → session.ready, and ready is just the WebSocket onConnect (rpc/session.ts:173), not the shell sync. Fifteen seconds for DNS + TLS + WS open is already generous. When it is exceeded the cause is almost always a dead relay or tunnel, and after this change every user on a dead T3 Connect link stares at "connecting" for 45s per attempt before the first retry, on web, desktop, and mobile. Your own description says the storm was what pushed healthy setups past the deadline, and #11811 removes the storm.

If there is still a reproducible case where a healthy connection needs more than 15s after #11811, please post it here with timings and we can talk about the number. Otherwise I'd close this once #11811 merges.

@Bil0000

Bil0000 commented Sep 14, 2026

Copy link
Copy Markdown
Contributor Author

Confirmed your review. Removed the 45-second timeout proposal in 0237d71 and restored the original 15-second deadline and tests. The slow-setup cases only injected delays; they were not evidence of a healthy connection needing more than 15 seconds.

I verified that #11811 contains the identical server patch and is merged. A merge preview of this corrected branch against main produces the same tree as main, so this PR is superseded and has nothing further to land.

Validation: 58 supervisor/RPC session tests plus 8 server-config tests pass. Client-runtime typecheck, scoped lint, formatting, and Ponytail review pass.

@Bil0000 Bil0000 closed this Sep 14, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XS 0-9 changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants