Repository navigation
[Bug]: Windows snapshot/evaluate timeouts evict the shared server-browser host and abort another open #16921
Description
Activity
Note
Grok responding on behalf of Julius.
Triage
Confirmed on
main(0678e4e23d86): one timed-out call on the server's own browser evicts the wholeserver-browserhost, which fails every other in-flight preview call in the environment, including an unrelatedpreview_open. The first snapshot stall itself is not explained here. It looks like #16567.What the code does
- A timeout on any call evicts the shared host. When a call gets no answer before its deadline, the broker calls
disconnect(..., true)(PreviewAutomationBroker.ts#L654-L665). Only the background metadata read is exempt.removeConnectionFromStatedrops every pending request and every session's host assignment on that connection (#L140-L162), andcloseConnectionfails each pending request withPreviewAutomationClientDisconnectedError(#L359-L379). That error's message is "Preview automation client server-browser disconnected during open." (previewAutomation.ts#L1060). - The host is one connection shared by the whole environment. The server browser registers once, with a fixed
clientIdandpreferred: true(ServerBrowser.ts#L2150-L2181), and preferred hosts are picked first (PreviewAutomationBroker.ts#L548-L556). So one tab's 15s snapshot timeout fails other tabs' and other threads' requests that are still in flight. That includes anopenon this host, whose 660s budget for a first-time Chromium install (#L562-L565) doesn't protect it from someone else's 15s deadline. Until the host loop registers again afterHOST_RECONNECT_DELAY(1s, ServerBrowser.ts#L88), new calls getPreviewAutomationNoAvailableHostError, whose message says "Do not retry" (previewAutomation.ts#L898). - Eviction doesn't stop the server-side work.
handleRequestwrapsrunOperationinEffect.tryPromisewith no abort signal and runs in a forked fiber (ServerBrowser.ts#L1706-L1731, #L2162-L2170). Its later reply is thrown away because the request is no longer pending under that connection (PreviewAutomationBroker.ts#L480-L493). For an evictedopen, the tab can still be created on the server after the caller has already been told the call failed. - A stalled snapshot keeps holding its tab. Snapshot and evaluate run in the tab's
SessionControlqueue (ServerBrowser.ts#L1605-L1617), and that queue has no per-action timeout (SessionControl.ts#L37-L44).evaluatesendsRuntime.evaluatewithout a timeout (ServerBrowserPage.ts#L413-L418). A later evaluate on the same tab waits behind the stuck snapshot, which fits the 17:37:18 evaluate timeout that followed the fresh tab's snapshot timeout.
The reporter's two hypotheses
includeImage: falsestill takes a screenshot: confirmed. The MCP handler removesincludeImageandsavebefore forwarding (handlers.ts#L211-L215).ServerBrowserPage.snapshotalways runspage.evaluate,ariaSnapshotandcaptureViewporttogether (ServerBrowserPage.ts#L191-L200), andcaptureViewportsends a barePage.captureScreenshot(#L170-L174).- Equal deadlines: confirmed, but there's no real race.
ariaSnapshotusesDEFAULT_TIMEOUT_MS = 15_000(ServerBrowserPage.ts#L66), andsnapshotgets the broker default of 15000ms. The broker's clock starts first, before the request is handed off and before any wait in the tab queue. So the broker always fires first, andariaSnapshot's own timeout can never get back to the caller. Thepage.evaluateand screenshot branches have no timeout at all.
Not confirmed
- Why the first snapshot on
example.comstalled on Windows. preview_snapshot on a background tab never returns and blocks every later action on that tab #16567 shows a desktop-rendered background tab whosePage.captureScreenshotnever settles. The report doesn't say whether the stalled tab was desktop-rendered or headless (the later working tab was headless), so this may or may not be the same trigger.
Fix direction
- Don't evict the in-process
server-browserhost when a call times out. It isn't unreachable, so fail only the call that timed out and leave the connection, the other pending requests and the session assignments alone. - Bound each snapshot stage (especially
Page.captureScreenshot) andRuntime.evaluatebelow the broker deadline, and release the tab'sSessionControlslot when the broker gives up. This overlaps with preview_snapshot on a background tab never returns and blocks every later action on that tab #16567. - Optionally, skip the screenshot when
includeImage: false.
Related
- preview_snapshot on a background tab never returns and blocks every later action on that tab #16567 (open): the same unbounded snapshot and stuck tab queue, on macOS background tabs. Fixing it would remove this trigger but wouldn't stop a timeout from evicting the host.
- [Bug]: A timed-out preview_evaluate keeps running in the renderer and makes all preview tools report "No preview automation host" until it finishes #16264, [Bug]: preview_resize outlives the broker deadline and evicts a live desktop host #16764, [Bug]: preview_wait_for still evicts a live host when its miss reply lands at the broker deadline #12898, [Bug]: Optional 500 ms preview metadata timeout disconnects the automation host #12273, [Bug]: preview_resize always times out (and evicts the automation host) when the main window is zoomed, because the guest webview reports innerWidth × hostZoom #12319, [Bug]: Preview loading bar never clears after a cross-origin iframe loads post-load, and automation timeouts drop the host #16779 (open): other timeouts that evict the host, mostly on the desktop host. This report adds the in-process server host and the in-flight
openfailure. - fix(web): the desktop browser host answers every request that arrives together, not only the last #15965 and fix(preview): reclaim hidden tabs after automation timeouts #14084 (open PRs) only change the desktop web host (
apps/web). The server browser consumes requests directly throughStream.runForEach, so neither covers this. fix(web): host local server browser tabs outside chat views #16620 (open PR) changes which engine renders a server tab, not the eviction or snapshot bounds. fix(server): keep preview hosts that report a waitFor miss at the deadline #12899 and fix(server): keep preview hosts when optional page metadata times out #12279 exempt onlywaitFormisses and metadata reads. - The Cloudflare challenge failures are out of scope here and are tracked in [Bug]: Environment browser stops at OpenAI Cloudflare verification while Edge works on the same Mac #16596 (with [Bug]: Cloudflare Turnstile challenge loops indefinitely in the built-in browser preview #5002 closed).
- A timeout on any call evicts the shared host. When a call gets no answer before its deadline, the broker calls
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Oct 7, 2026 Submitted #16941 on top of your merged child-frame fix #16939. It retains the in-process host on per-request timeouts, binds execution and queued cancellation to the original request lifetime, bounds browser reads, and skips capture for text-only snapshots unless saving a PNG.
181 focused tests pass, including real Chromium tests. The additional Windows check uses Electron 44.4.2 / Chromium 152.0.7977.130 with actual T3 browser code and the merged relay: text-only inspection/clicks work on a hidden guest, and timed-out capture/evaluation leaves the same tab usable. Independent review identified two deadline/cancellation gaps that were fixed before submission; the cancelled-click case was reproduced before its fix.
The PR documents the remaining limit: a fully hidden native guest can still fail PNG capture, but the timeout is bounded. It does not duplicate the child-frame fix or claim a verified Cloudflare login on the original account.
Same host-eviction chain on macOS, so this isn't specific to Windows.
Environment: T3 Code Nightly
0.0.46-nightly.20261008.2819(5e22256), macOS 27.0 arm64, desktop app with a local server, Claude Code and Codex threads.Server trace, 2026-10-08 (local time):
09:49:20 PreviewAutomationBroker.invoke PreviewAutomationTimeoutError: Preview automation evaluate timed out after 15000ms. 09:49:35 PreviewAutomationBroker.disconnect (x2) 09:49:36 PreviewAutomationBroker.connect / focusHost 09:49:41 PreviewAutomationBroker.invoke PreviewAutomationTimeoutError: Preview automation navigate timed out after 15000ms. 09:49:56 PreviewAutomationBroker.disconnect (x2) 09:49:57 PreviewAutomationBroker.connect / focusHostEach time,
server-browserwas the only registered host. After each eviction the agent's lease and tab assignment were gone. The agent then opened a fresh tab withpreview_open. When the desktop panel didn't attach it withinDESKTOP_ATTACH_TIMEOUT, that tab ran headless in an isolated context (#16901). The user sees this as "the browser was signed in, then the agent says it isn't and loses the connection." It happens often enough that they have to step in during most browser tasks.A related way the same session is lost: a nightly auto-update install (
desktop.updates.installat 09:06:19) restarted the backend. The agent's next call failed withNo server preview tab is open for this thread. Call preview_open first.PR #16941 (keep the host on per-request timeouts) looks like it covers the main path here.
Posted by Claude Code (Claude Opus 5.5) via
t3 triage, on behalf of the user.@jamielmccormick Your trace matches the timeout/host-eviction path covered by #16941. I added an explicit navigation-timeout case in f07b86a: the host stays registered, both sessions keep their tab assignments, and the unrelated pending open can complete. All 58 broker tests pass, along with the server typecheck and targeted lint.
This is shared broker coverage; I have not run a native macOS desktop test. The backend-restart case is separate: this PR preserves registrations and tab assignments during per-request timeouts, but does not persist them across a backend restart.
Reacted by Jamie McCormickStill happens on macOS 26.7.1 arm64, Nightly
0.0.46-nightly.20261008.2833, which is newer than the .2819 report above. Times are UTC.- In one hour of
server.trace(2026-10-09, 08:26–09:31), every broker timeout was followed byPreviewAutomationBroker.disconnect/closeConnectionand aconnectabout 1 s later. That happened 23 times, after snapshot, evaluate, click, waitFor and navigate timeouts. - Provider logs from 6–8 Oct show 43 "No preview automation host is available" errors. 42 of them came within 20 s of a timeout in the same thread.
- Codex threads hit it hardest because they send several preview calls at once. One timeout at 06:10:36 on 8 Oct took 7 parallel calls down with it.
- After the reconnect,
preview_statusoften returnedavailable: false, tabId: null, and the agent had to open a new tab, which came up signed out.
- In one hour of
Before submitting
.2735ordinary-page evidence and a trace of one timeout aborting a separateopenrequest. I am not claiming a newly established root cause for every symptom; please consolidate with an existing report if appropriate.Area
apps/server/ preview automation broker and server browser, with the Windows desktop/native-rendering path also involved.Summary (original 7 October report)
On Windows 11, T3 browser snapshots and evaluations timed out after 15 seconds on
example.com, even after recreating the tab and downgrading from.2774to exact.2735. The timeout disconnected the shared automation host; the following click/status calls returned "No preview automation host is available".During a later investigation, a snapshot already in flight timed out and disconnected
server-browser, causing a differentopenrequest to fail after only about 6.5 seconds. The host registered again roughly one second later. This cascade is visible in the local traces and matches the broker implementation.The original snapshot stall remains unexplained. The browser is intermittently usable: a later fresh
.2735tab, verified as Windows Chrome headless shell 154, successfully completed DOM reads, a snapshot, and a link click. Cloudflare verification remained blocked in that engine. Earlier Windows Cloudflare failure used native Electron/Chromium 152, so these should not be treated as a single proven engine-specific defect.Latest live retest — 10 October 2026
Tested the unpatched official
0.0.46-nightly.20261010.2922(bd2346eda). PR #16941 is still unmerged and was not installed. This tab actually ran on Windows Chrome headless shell154.0.8037.92(Chrome/154, Win32,navigator.webdriver === false), not the app's Electron44.4.5/ embedded Chromium152.0.7977.130renderer.There was a problem with verification. Please reload and try again.with Sign in disabled.Just a moment.../Performing security verificationinstead, with a human-verification checkbox visible in the saved screenshot.The page changed while capturing its snapshot. Take another snapshot.. Evaluation and cookie-banner dismissal still worked. A challenge-page PNG succeeded after reload.The original ordinary-page 15-second snapshot failure and missing-host cascade were not reproduced in this retest. Current upstream's evaluation termination and response grace are present in this release. These successful normal typed-timeout paths do not establish recovery from a genuinely unanswered read/capture stage beyond the grace period or cancellation of expired queued/active work; those remain covered by PR #16941's code and regression tests.
Cloudflare verification remains unresolved in this fresh headless session. Initial 403 responses were observed at
/cdn-cgi/zaraz/t; their relationship to verification failure is unknown, and the earlier origin-warning cause was not re-established. The page-change snapshot errors did not turn into a 15-second hang or host eviction. No challenge was solved, no credentials entered, and no sign-in attempted. Manual Take control, native Electron hidden capture and native macOS were not retested. The original evidence below remains historical evidence, not a claim that every symptom persists on nightly 2922.Steps to reproduce
This is the recorded intermittent sequence; a later fresh headless tab succeeded.
.2735, callpreview_status, thenpreview_openand navigate tohttps://example.com.Example Domainandloading: false.preview_snapshot({tabId, includeImage:false}). The recorded call times out after 15,000 ms.t3_preview_closeand open a fresh tab at the same URL withreuseExistingTab:false.document.titleand the first link succeeds once.The later cascade was captured while another snapshot was pending. A new
preview_openrequest overlapped it and failed when that snapshot's deadline disconnected the shared host. This was not anopenrequest reaching its own timeout.Expected behavior
Successful navigation should leave ordinary page inspection and interaction usable. A slow snapshot or evaluation should not take out unrelated requests on a healthy shared in-process browser host. If a host is recovering, the error should describe that temporary state accurately.
Preserve the existing protection against replaying an action whose effects are uncertain. No automatic replay of clicks or other mutations is requested.
Actual behavior and trace evidence
All times are UTC on 7 October 2026; Amsterdam local time was UTC+2. Individual environment, thread, session, tab, and trace identifiers have been omitted.
disconnect/closeConnectionopenrequest starts waitingdisconnect/closeConnection; the pendingopenfails after 6,484 ms withserver-browser disconnected during openThe Windows app processes remained running. Process existence alone would not establish browser health, but the disconnection/re-registration sequence above is independently visible in server traces.
Relevant exact errors:
{ "error": { "_tag": "PreviewAutomationTimeoutError", "operation": "snapshot", "failureCount": 1, "message": "Preview automation snapshot timed out after 15000ms." } }The subsequent click error, with the environment identifier replaced, was:
A corresponding
for statuserror followed. No click was performed in that failed test.Source findings and investigation limits
Release
.2735resolves tofd1c3386c4d60f3477ab3f13c87537848de099f5;.2774resolves to611132c171f3a821bd2e32f22261135cef6330ac.includeImage:falsedoes not bypass image capture. The MCP handler strips output-selection options. ServerBrowserPage.snapshot still waits concurrently for page evaluation, a boxed accessibility snapshot, andPage.captureScreenshot. The stalled branch has not been isolated. The accessibility timeout and broker deadline are both 15 seconds; whether a deadline race contributed here is unproven.runtime:serverdoes not identify the rendering engine. ServerBrowser's native-attachment path waits up to ten seconds for its desktop page and otherwise uses headless Chromium. The app's Electron version alone cannot identify a tab's engine.Possible maintainer directions are to bound/diagnose individual snapshot stages, cancel or retire the affected operation/tab without evicting unrelated healthy work, and distinguish a recovering host from a genuinely absent one. These are suggestions, not an implemented or verified fix.
The latest release available when checked was
.2787(f570bd21663f56ce94c41829d3b7d72886e25a34). Its broker, server browser, snapshot implementation, and CDP relay files were byte-identical to.2735; its Electron host component differed..2787was not installed or tested, so no claim is made that upgrading cannot help.Successful retest and remaining Cloudflare failure
At 19:45–19:47 UTC, without applying a source patch or another version change, a fresh Windows
.2735tab completed:textContentandinnerText;Learn more, reaching the IANA example-domains page.The tab reported
runtime:server; its page user agent was:navigator.webdriverwastrue. The Windows Chrome headless shell process and version were independently checked. This establishes recovery of basic automation on that engine; it does not identify why the original tab stalled or prove native Electron recovery.Opening a fresh
https://dash.cloudflare.com/logintab at 19:47:54 UTC still reachedJust a moment.... At 19:48:24 it showed security verification,readyState:complete, no Sign in button, a 403 resource error, andNo available adapters.warnings. At 19:51:42 it was still at security verification without a Sign in button. Snapshot inspection remained usable. No challenge was solved, credentials entered, identity altered, or session data transferred.Cloudflare's supported-browser documentation describes limited support for embedded browsers and excludes automated browsers from production challenge solving. This provides compatibility context for the headless failure, not a diagnosis of the earlier native Electron failure.
Impact
Major degradation: the Cloudflare email-forwarding workflow could not be completed through T3's browser, and ordinary page automation also failed intermittently. A practical fallback eventually worked in an existing signed-in external Chrome session.
Version and environment
10.0.26200, build2620031.0.21924.61B450M-HDV R4.044.4.2, Chromium152.0.7977.130, Node24.21.0154.0.8037.9224.04.5 LTSx86_64; kernel7.0.0-31-generic; Ryzen AI MAX+ 395 / Radeon 8060S; 32 logical CPUsEarlier symptoms and all attempted workarounds
Initial Linux sandbox problem, native Windows Cloudflare failure, control tests, downgrade, and external Chrome errors
Ubuntu
.2752: At 12:10 and 12:15 UTC, browser open failed with the explicit Ubuntu AppArmor sandbox message. Running the official version-specifict3 browser setupwith sudo fixed startup. Setup allowed/etc/apparmor.d/t3-chrome-headless-shellwith user-namespace permission; the global restriction remained enabled, and missing libraries were not found. Subsequent open and snapshot succeeded. This startup problem is resolved.Cloudflare's
Verify you are humanstill did not complete when I used Take control and clicked manually. After releasing control, status reported the agent as owner, but snapshot/evaluation timed out after 15 seconds. Closing the stuck tab and opening a fresh one restored ordinary reads and agent interaction temporarily.A first button test injected into a blank tab was readable/clickable by tools, but the visible pane stayed on T3's start screen because the tab had no actual URL. This was a test-setup mistake. Serving a real localhost HTTP button page corrected it, and I confirmed manual clicking worked. Cloudflare verification still failed afterward. Agent calls while I retained control were correctly rejected; later timeouts also occurred after control returned to the agent.
One subsequent Linux close attempt returned
OrchestratorMcpFailure, codeorchestration_error, messageThe operation could not be completed.A fresh open on the official Cloudflare login URL succeeded but remained at the challenge.Windows
.2774: At 12:57 UTC, the Cloudflare login capture logged a generic 403 and 19 repetitions of:The visible error was
There was a problem with verification. Please reload and try again.Reloading by navigating to/logindid not resolve it. DOM inspection confirmedSign inwas visible and disabled.navigator.webdriverwasfalse, and the user agent includedT3Code(Nightly)/0.0.46-nightly.20261007.2774 Chrome/152.0.7977.130 Electron/44.4.2. This was native Electron, not the later headless-154 retest.The snapshot's
interactiveElementslist and a later iframe query were empty. The generic 403 was not correlated to a particular request, and the origin warning's source/frame was not traced. None establishes the exact verification cause or an iframe-removal bug.Downgrade: Windows was replaced/restarted on exact
.2735; installer signature, official asset checksums, executable/install version, running app, and T3 environment version were checked. The active Linux backend and installed AppImage link were changed to.2735, with backups and endpoint health checks. Older unrelated Linux processes were outside the active backend route and were not all restarted. A fresh Linux preview after the downgrade was not tested. HTTP backend reachability is not browser-health evidence. The Windows ordinary-page failures described above persisted; they do not prove the identical native Cloudflare console failure on.2735.External Chrome: The offered old
t3_fastpc.browser_tabsroute failed withMcpServerError: Tool browser_tabs not found(INVALID_ARGUMENT, JSON-RPC-32602). The supported bundledchrome@openai-bundledplugin26.908.40834, bootstrapped through documentedsetupBrowserRuntime()innode_repl, worked. Existing signed-in Cloudflare tabs were readable/clickable and the email-forwarding task completed.In a later Chrome task, the Windows computer-use route returned
codex app-server exited before returning response 1three times. Twotab.screenshot(...)attempts, including a smaller clip, failed withTimed out after 5000ms waiting for CDP command Page.captureScreenshot.The same Chrome session remained reachable;tab.ax.get('screenshot')produced usable screenshots. These are separate route-specific observations, not proof of a shared T3 preview cause. A script-variable error and a separate approval-layer rejection were also excluded from browser-failure diagnoses.The external Chrome comparison reused a logged-in profile; a controlled fresh challenge test with matching profile/cookies/network was not performed. It is a practical workaround rather than proof of a native Electron fix.
Related reports and proposed fixes
.2735loading/iframe issues and a timeout cascade on macOS. No injected iframe or stuck loading-bar trigger was established in this Example Domain test..0.0.42missing-host report. Its metadata-timeout proposal fix(server): keep the preview host after an optional status timeout #14552 does not explain the recorded primary 15-second snapshot/evaluation deadlines;.2735already avoids eviction for its optional metadata read..2774failure above already retained the native Electron user agent..2735preferred server browser consumes requests directly throughStream.runForEach; its applicability to these failures is unproven.There is no verified complete fix from this investigation. Resolved sandbox startup, intermittent basic-automation recovery, a working external Chrome fallback, and upstream candidate fixes should be distinguished from a repaired T3 Cloudflare workflow.