Skip to content

[Bug]: Windows snapshot/evaluate timeouts evict the shared server-browser host and abort another open #16921

Description

@hogeheer499-commits

Before submitting

Area

apps/server / preview automation broker and server browser, with the Windows desktop/native-rendering path also involved.

Summary (original 7 October report)

On Windows 11, T3 browser snapshots and evaluations timed out after 15 seconds on example.com, even after recreating the tab and downgrading from .2774 to exact .2735. The timeout disconnected the shared automation host; the following click/status calls returned "No preview automation host is available".

During a later investigation, a snapshot already in flight timed out and disconnected server-browser, causing a different open request to fail after only about 6.5 seconds. The host registered again roughly one second later. This cascade is visible in the local traces and matches the broker implementation.

The original snapshot stall remains unexplained. The browser is intermittently usable: a later fresh .2735 tab, verified as Windows Chrome headless shell 154, successfully completed DOM reads, a snapshot, and a link click. Cloudflare verification remained blocked in that engine. Earlier Windows Cloudflare failure used native Electron/Chromium 152, so these should not be treated as a single proven engine-specific defect.

Latest live retest — 10 October 2026

Tested the unpatched official 0.0.46-nightly.20261010.2922 (bd2346eda). PR #16941 is still unmerged and was not installed. This tab actually ran on Windows Chrome headless shell 154.0.8037.92 (Chrome/154, Win32, navigator.webdriver === false), not the app's Electron 44.4.5 / embedded Chromium 152.0.7977.130 renderer.

Check Result
Ordinary open/navigation, DOM evaluation and text snapshots Passed on example.com and an isolated local test page.
Snapshot-ref clicking, typing and background PNG save Passed; counter and typed text verified, saved image visually inspected.
Deliberate 300 ms navigation timeout Expected timeout after 346 ms; host and same-tab navigation/snapshot recovered.
Unresolved evaluation Expected timeout after 15,055 ms; same-tab evaluation recovered in 46 ms.
Concurrent other-tab request A 25-second wait completed after 17,116 ms while the first tab's evaluation timed out; no host-loss error. Later snapshot/ref click also worked.
Fresh incognito Cloudflare login Still showed There was a problem with verification. Please reload and try again. with Sign in disabled.
Normal Cloudflare reload Reached Just a moment... / Performing security verification instead, with a human-verification checkbox visible in the saved screenshot.
Cloudflare snapshots Three settled-login attempts returned The page changed while capturing its snapshot. Take another snapshot.. Evaluation and cookie-banner dismissal still worked. A challenge-page PNG succeeded after reload.

The original ordinary-page 15-second snapshot failure and missing-host cascade were not reproduced in this retest. Current upstream's evaluation termination and response grace are present in this release. These successful normal typed-timeout paths do not establish recovery from a genuinely unanswered read/capture stage beyond the grace period or cancellation of expired queued/active work; those remain covered by PR #16941's code and regression tests.

Cloudflare verification remains unresolved in this fresh headless session. Initial 403 responses were observed at /cdn-cgi/zaraz/t; their relationship to verification failure is unknown, and the earlier origin-warning cause was not re-established. The page-change snapshot errors did not turn into a 15-second hang or host eviction. No challenge was solved, no credentials entered, and no sign-in attempted. Manual Take control, native Electron hidden capture and native macOS were not retested. The original evidence below remains historical evidence, not a claim that every symptom persists on nightly 2922.

Steps to reproduce

This is the recorded intermittent sequence; a later fresh headless tab succeeded.

  1. On Windows T3 .2735, call preview_status, then preview_open and navigate to https://example.com.
  2. Navigation succeeds, with title Example Domain and loading: false.
  3. Call preview_snapshot({tabId, includeImage:false}). The recorded call times out after 15,000 ms.
  4. Close that tab through t3_preview_close and open a fresh tab at the same URL with reuseExistingTab:false.
  5. A small evaluation of document.title and the first link succeeds once.
  6. Another snapshot times out after 15,000 ms. Evaluating body text and links then times out too.
  7. The next click on the IANA link is rejected with the missing-host error. An immediate status call is rejected too.

The later cascade was captured while another snapshot was pending. A new preview_open request overlapped it and failed when that snapshot's deadline disconnected the shared host. This was not an open request reaching its own timeout.

Expected behavior

Successful navigation should leave ordinary page inspection and interaction usable. A slow snapshot or evaluation should not take out unrelated requests on a healthy shared in-process browser host. If a host is recovering, the error should describe that temporary state accurately.

Preserve the existing protection against replaying an action whose effects are uncertain. No automatic replay of clicks or other mutations is requested.

Actual behavior and trace evidence

All times are UTC on 7 October 2026; Amsterdam local time was UTC+2. Individual environment, thread, session, tab, and trace identifiers have been omitted.

Time UTC Recorded event
17:33:56 Navigation to Example Domain succeeds
17:34:12.208–17:34:27.214 Broker waits about 15,005 ms for the first snapshot, then returns a timeout
17:34:27.213 disconnect / closeConnection
17:34:28.223 Host connects again
17:36:27 Old tab closed and fresh tab opened
17:36:41 Evaluation returns the title and IANA link
17:36:49.413–17:37:04.425 Fresh-tab snapshot times out after about 15,012 ms
17:37:04.425 Host disconnected again
17:37:05.432 Host connects again
17:37:18.315–17:37:33.316 Evaluation times out after about 15,001 ms
17:37:33.316 Host disconnected; following click/status cannot find a host
17:37:34.327 Host connects again
17:50:19 Later status reports the tab available; complete recovery was not tested at that point
19:33:14.824 Another snapshot starts waiting for a broker response
19:33:23.342 A separate open request starts waiting
19:33:29.826 disconnect / closeConnection; the pending open fails after 6,484 ms with server-browser disconnected during open
19:33:29.842 The snapshot's failed span finishes at about 15,018 ms
19:33:30.852 Host connects again

The Windows app processes remained running. Process existence alone would not establish browser health, but the disconnection/re-registration sequence above is independently visible in server traces.

Relevant exact errors:

Preview automation snapshot timed out after 15000ms.
Preview automation evaluate timed out after 15000ms.
Preview automation client server-browser disconnected during open.
{
  "error": {
    "_tag": "PreviewAutomationTimeoutError",
    "operation": "snapshot",
    "failureCount": 1,
    "message": "Preview automation snapshot timed out after 15000ms."
  }
}

The subsequent click error, with the environment identifier replaced, was:

No preview automation host is available for click in environment [ENVIRONMENT_ID]. Preview tools run in a T3 Code desktop app that is open and connected to this environment; a headless server has no browser of its own. Do not retry. To check a page, use a headless browser from the shell, such as Playwright, or curl, or ask the user to open this thread in the T3 Code desktop app.

A corresponding for status error followed. No click was performed in that failed test.

Source findings and investigation limits

Release .2735 resolves to fd1c3386c4d60f3477ab3f13c87537848de099f5; .2774 resolves to 611132c171f3a821bd2e32f22261135cef6330ac.

  1. Why other requests get disconnected is established. In PreviewAutomationBroker.ts, an unanswered primary request disconnects its connection. closeConnection fails the other pending requests on that connection. This explains the captured cascade, not the initial stall.
  2. The existing recovery does run. ServerBrowser's host loop reconnects after one second; the traces agree. fix(preview): recover host registration after request timeouts #12535 is already merged. This is not simply the older permanently-unregistered-host problem.
  3. includeImage:false does not bypass image capture. The MCP handler strips output-selection options. ServerBrowserPage.snapshot still waits concurrently for page evaluation, a boxed accessibility snapshot, and Page.captureScreenshot. The stalled branch has not been isolated. The accessibility timeout and broker deadline are both 15 seconds; whether a deadline race contributed here is unproven.
  4. runtime:server does not identify the rendering engine. ServerBrowser's native-attachment path waits up to ten seconds for its desktop page and otherwise uses headless Chromium. The app's Electron version alone cannot identify a tab's engine.

Possible maintainer directions are to bound/diagnose individual snapshot stages, cancel or retire the affected operation/tab without evicting unrelated healthy work, and distinguish a recovering host from a genuinely absent one. These are suggestions, not an implemented or verified fix.

The latest release available when checked was .2787 (f570bd21663f56ce94c41829d3b7d72886e25a34). Its broker, server browser, snapshot implementation, and CDP relay files were byte-identical to .2735; its Electron host component differed. .2787 was not installed or tested, so no claim is made that upgrading cannot help.

Successful retest and remaining Cloudflare failure

At 19:45–19:47 UTC, without applying a source patch or another version change, a fresh Windows .2735 tab completed:

  • DOM reads, including body textContent and innerText;
  • a full snapshot, with accessibility refs and an internally captured PNG;
  • a click using the snapshot-provided ref for Learn more, reaching the IANA example-domains page.

The tab reported runtime:server; its page user agent was:

Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) HeadlessChrome/154.0.8037.92 Safari/537.36

navigator.webdriver was true. The Windows Chrome headless shell process and version were independently checked. This establishes recovery of basic automation on that engine; it does not identify why the original tab stalled or prove native Electron recovery.

Opening a fresh https://dash.cloudflare.com/login tab at 19:47:54 UTC still reached Just a moment.... At 19:48:24 it showed security verification, readyState:complete, no Sign in button, a 403 resource error, and No available adapters. warnings. At 19:51:42 it was still at security verification without a Sign in button. Snapshot inspection remained usable. No challenge was solved, credentials entered, identity altered, or session data transferred.

Cloudflare's supported-browser documentation describes limited support for embedded browsers and excludes automated browsers from production challenge solving. This provides compatibility context for the headless failure, not a diagnosis of the earlier native Electron failure.

Impact

Major degradation: the Cloudflare email-forwarding workflow could not be completed through T3's browser, and ordinary page automation also failed intermittently. A practical fallback eventually worked in an existing signed-in external Chrome session.

Version and environment

Component Verified details
Windows OS Windows 11 Pro x64; 10.0.26200, build 26200
CPU / RAM Ryzen 5 3600, 6 cores / 12 threads; 32 GiB RAM
GPU / driver Radeon RX550/550 Series; driver 31.0.21924.61
Motherboard B450M-HDV R4.0
Windows T3 app runtime Electron 44.4.2, Chromium 152.0.7977.130, Node 24.21.0
Later Windows server browser Chrome headless shell 154.0.8037.92
Agent Codex through T3 Code
Earlier remote Linux machine Ubuntu 24.04.5 LTS x86_64; kernel 7.0.0-31-generic; Ryzen AI MAX+ 395 / Radeon 8060S; 32 logical CPUs

Earlier symptoms and all attempted workarounds

Initial Linux sandbox problem, native Windows Cloudflare failure, control tests, downgrade, and external Chrome errors

Ubuntu .2752: At 12:10 and 12:15 UTC, browser open failed with the explicit Ubuntu AppArmor sandbox message. Running the official version-specific t3 browser setup with sudo fixed startup. Setup allowed /etc/apparmor.d/t3-chrome-headless-shell with user-namespace permission; the global restriction remained enabled, and missing libraries were not found. Subsequent open and snapshot succeeded. This startup problem is resolved.

Cloudflare's Verify you are human still did not complete when I used Take control and clicked manually. After releasing control, status reported the agent as owner, but snapshot/evaluation timed out after 15 seconds. Closing the stuck tab and opening a fresh one restored ordinary reads and agent interaction temporarily.

A first button test injected into a blank tab was readable/clickable by tools, but the visible pane stayed on T3's start screen because the tab had no actual URL. This was a test-setup mistake. Serving a real localhost HTTP button page corrected it, and I confirmed manual clicking worked. Cloudflare verification still failed afterward. Agent calls while I retained control were correctly rejected; later timeouts also occurred after control returned to the agent.

One subsequent Linux close attempt returned OrchestratorMcpFailure, code orchestration_error, message The operation could not be completed. A fresh open on the official Cloudflare login URL succeeded but remained at the challenge.

Windows .2774: At 12:57 UTC, the Cloudflare login capture logged a generic 403 and 19 repetitions of:

Failed to execute 'postMessage' on 'DOMWindow': The target origin provided ('https://challenges.cloudflare.com') does not match the recipient window's origin ('https://dash.cloudflare.com').

The visible error was There was a problem with verification. Please reload and try again. Reloading by navigating to /login did not resolve it. DOM inspection confirmed Sign in was visible and disabled. navigator.webdriver was false, and the user agent included T3Code(Nightly)/0.0.46-nightly.20261007.2774 Chrome/152.0.7977.130 Electron/44.4.2. This was native Electron, not the later headless-154 retest.

The snapshot's interactiveElements list and a later iframe query were empty. The generic 403 was not correlated to a particular request, and the origin warning's source/frame was not traced. None establishes the exact verification cause or an iframe-removal bug.

Downgrade: Windows was replaced/restarted on exact .2735; installer signature, official asset checksums, executable/install version, running app, and T3 environment version were checked. The active Linux backend and installed AppImage link were changed to .2735, with backups and endpoint health checks. Older unrelated Linux processes were outside the active backend route and were not all restarted. A fresh Linux preview after the downgrade was not tested. HTTP backend reachability is not browser-health evidence. The Windows ordinary-page failures described above persisted; they do not prove the identical native Cloudflare console failure on .2735.

External Chrome: The offered old t3_fastpc.browser_tabs route failed with McpServerError: Tool browser_tabs not found (INVALID_ARGUMENT, JSON-RPC -32602). The supported bundled chrome@openai-bundled plugin 26.908.40834, bootstrapped through documented setupBrowserRuntime() in node_repl, worked. Existing signed-in Cloudflare tabs were readable/clickable and the email-forwarding task completed.

In a later Chrome task, the Windows computer-use route returned codex app-server exited before returning response 1 three times. Two tab.screenshot(...) attempts, including a smaller clip, failed with Timed out after 5000ms waiting for CDP command Page.captureScreenshot. The same Chrome session remained reachable; tab.ax.get('screenshot') produced usable screenshots. These are separate route-specific observations, not proof of a shared T3 preview cause. A script-variable error and a separate approval-layer rejection were also excluded from browser-failure diagnoses.

The external Chrome comparison reused a logged-in profile; a controlled fresh challenge test with matching profile/cookies/network was not performed. It is a practical workaround rather than proof of a native Electron fix.

Related reports and proposed fixes

There is no verified complete fix from this investigation. Resolved sandbox startup, intermittent basic-automation recovery, a working external Chrome fallback, and upstream candidate fixes should be distinguished from a repaired T3 Cloudflare workflow.

Activity

  1. juliusmarminge commented on Oct 7, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Confirmed on main (0678e4e23d86): one timed-out call on the server's own browser evicts the whole server-browser host, which fails every other in-flight preview call in the environment, including an unrelated preview_open. The first snapshot stall itself is not explained here. It looks like #16567.

    What the code does

    1. A timeout on any call evicts the shared host. When a call gets no answer before its deadline, the broker calls disconnect(..., true) (PreviewAutomationBroker.ts#L654-L665). Only the background metadata read is exempt. removeConnectionFromState drops every pending request and every session's host assignment on that connection (#L140-L162), and closeConnection fails each pending request with PreviewAutomationClientDisconnectedError (#L359-L379). That error's message is "Preview automation client server-browser disconnected during open." (previewAutomation.ts#L1060).
    2. The host is one connection shared by the whole environment. The server browser registers once, with a fixed clientId and preferred: true (ServerBrowser.ts#L2150-L2181), and preferred hosts are picked first (PreviewAutomationBroker.ts#L548-L556). So one tab's 15s snapshot timeout fails other tabs' and other threads' requests that are still in flight. That includes an open on this host, whose 660s budget for a first-time Chromium install (#L562-L565) doesn't protect it from someone else's 15s deadline. Until the host loop registers again after HOST_RECONNECT_DELAY (1s, ServerBrowser.ts#L88), new calls get PreviewAutomationNoAvailableHostError, whose message says "Do not retry" (previewAutomation.ts#L898).
    3. Eviction doesn't stop the server-side work. handleRequest wraps runOperation in Effect.tryPromise with no abort signal and runs in a forked fiber (ServerBrowser.ts#L1706-L1731, #L2162-L2170). Its later reply is thrown away because the request is no longer pending under that connection (PreviewAutomationBroker.ts#L480-L493). For an evicted open, the tab can still be created on the server after the caller has already been told the call failed.
    4. A stalled snapshot keeps holding its tab. Snapshot and evaluate run in the tab's SessionControl queue (ServerBrowser.ts#L1605-L1617), and that queue has no per-action timeout (SessionControl.ts#L37-L44). evaluate sends Runtime.evaluate without a timeout (ServerBrowserPage.ts#L413-L418). A later evaluate on the same tab waits behind the stuck snapshot, which fits the 17:37:18 evaluate timeout that followed the fresh tab's snapshot timeout.

    The reporter's two hypotheses

    • includeImage: false still takes a screenshot: confirmed. The MCP handler removes includeImage and save before forwarding (handlers.ts#L211-L215). ServerBrowserPage.snapshot always runs page.evaluate, ariaSnapshot and captureViewport together (ServerBrowserPage.ts#L191-L200), and captureViewport sends a bare Page.captureScreenshot (#L170-L174).
    • Equal deadlines: confirmed, but there's no real race. ariaSnapshot uses DEFAULT_TIMEOUT_MS = 15_000 (ServerBrowserPage.ts#L66), and snapshot gets the broker default of 15000ms. The broker's clock starts first, before the request is handed off and before any wait in the tab queue. So the broker always fires first, and ariaSnapshot's own timeout can never get back to the caller. The page.evaluate and screenshot branches have no timeout at all.

    Not confirmed

    Fix direction

    • Don't evict the in-process server-browser host when a call times out. It isn't unreachable, so fail only the call that timed out and leave the connection, the other pending requests and the session assignments alone.
    • Bound each snapshot stage (especially Page.captureScreenshot) and Runtime.evaluate below the broker deadline, and release the tab's SessionControl slot when the broker gives up. This overlaps with preview_snapshot on a background tab never returns and blocks every later action on that tab #16567.
    • Optionally, skip the screenshot when includeImage: false.

    Related

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 7, 2026
  3. hogeheer499-commits commented on Oct 7, 2026

    @hogeheer499-commits
    Author

    Submitted #16941 on top of your merged child-frame fix #16939. It retains the in-process host on per-request timeouts, binds execution and queued cancellation to the original request lifetime, bounds browser reads, and skips capture for text-only snapshots unless saving a PNG.

    181 focused tests pass, including real Chromium tests. The additional Windows check uses Electron 44.4.2 / Chromium 152.0.7977.130 with actual T3 browser code and the merged relay: text-only inspection/clicks work on a hidden guest, and timed-out capture/evaluation leaves the same tab usable. Independent review identified two deadline/cancellation gaps that were fixed before submission; the cancelled-click case was reproduced before its fix.

    The PR documents the remaining limit: a fully hidden native guest can still fail PNG capture, but the timeout is bounded. It does not duplicate the child-frame fix or claim a verified Cloudflare login on the original account.

  4. jamielmccormick commented on Oct 8, 2026

    @jamielmccormick

    Same host-eviction chain on macOS, so this isn't specific to Windows.

    Environment: T3 Code Nightly 0.0.46-nightly.20261008.2819 (5e22256), macOS 27.0 arm64, desktop app with a local server, Claude Code and Codex threads.

    Server trace, 2026-10-08 (local time):

    09:49:20 PreviewAutomationBroker.invoke  PreviewAutomationTimeoutError: Preview automation evaluate timed out after 15000ms.
    09:49:35 PreviewAutomationBroker.disconnect (x2)
    09:49:36 PreviewAutomationBroker.connect / focusHost
    09:49:41 PreviewAutomationBroker.invoke  PreviewAutomationTimeoutError: Preview automation navigate timed out after 15000ms.
    09:49:56 PreviewAutomationBroker.disconnect (x2)
    09:49:57 PreviewAutomationBroker.connect / focusHost
    

    Each time, server-browser was the only registered host. After each eviction the agent's lease and tab assignment were gone. The agent then opened a fresh tab with preview_open. When the desktop panel didn't attach it within DESKTOP_ATTACH_TIMEOUT, that tab ran headless in an isolated context (#16901). The user sees this as "the browser was signed in, then the agent says it isn't and loses the connection." It happens often enough that they have to step in during most browser tasks.

    A related way the same session is lost: a nightly auto-update install (desktop.updates.install at 09:06:19) restarted the backend. The agent's next call failed with No server preview tab is open for this thread. Call preview_open first.

    PR #16941 (keep the host on per-request timeouts) looks like it covers the main path here.

    Posted by Claude Code (Claude Opus 5.5) via t3 triage, on behalf of the user.

  5. hogeheer499-commits commented on Oct 8, 2026

    @hogeheer499-commits
    Author

    @jamielmccormick Your trace matches the timeout/host-eviction path covered by #16941. I added an explicit navigation-timeout case in f07b86a: the host stays registered, both sessions keep their tab assignments, and the unrelated pending open can complete. All 58 broker tests pass, along with the server typecheck and targeted lint.

    This is shared broker coverage; I have not run a native macOS desktop test. The backend-restart case is separate: this PR preserves registrations and tab assignments during per-request timeouts, but does not persist them across a backend restart.

  6. marius-ciocoiu commented on Oct 9, 2026

    @marius-ciocoiu

    Still happens on macOS 26.7.1 arm64, Nightly 0.0.46-nightly.20261008.2833, which is newer than the .2819 report above. Times are UTC.

    • In one hour of server.trace (2026-10-09, 08:26–09:31), every broker timeout was followed by PreviewAutomationBroker.disconnect/closeConnection and a connect about 1 s later. That happened 23 times, after snapshot, evaluate, click, waitFor and navigate timeouts.
    • Provider logs from 6–8 Oct show 43 "No preview automation host is available" errors. 42 of them came within 20 s of a timeout in the same thread.
    • Codex threads hit it hardest because they send several preview calls at once. One timeout at 06:10:36 on 8 Oct took 7 parallel calls down with it.
    • After the reconnect, preview_status often returned available: false, tabId: null, and the agent had to open a new tab, which came up signed out.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions