Skip to content

[Bug]: providerHealthRefreshInterval has no floor tied to HEALTH_CHECK_TIMEOUT, so provider health probes overlap concurrently #12000

Description

@indishere

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. On Windows, enable the managed Antigravity provider.
  2. Open Settings → Background activity, choose a custom profile, and set Provider health refresh interval to 30 seconds (the UI permits this; performance is 1 min and the balanced default is 5 min).
  3. Leave T3 Code running and watch agy_acp_server.exe in Task Manager, or count _MEI* directories in %LOCALAPPDATA%\Temp.

Expected behavior

A process-based provider health probe should not be re-issued while the previous probe for that same provider is still in flight. The refresh interval should be clamped to at least the probe timeout, or probes should be de-duplicated per provider so at most one runs at a time.

Actual behavior

providerHealthRefreshInterval and HEALTH_CHECK_TIMEOUT are configured independently, with no floor relating them:

Setting Value
HEALTH_CHECK_TIMEOUT "90 seconds"
battery-saver preset 15 min
DEFAULT_PROVIDER_HEALTH_REFRESH_INTERVAL (balanced) 5 min
performance preset 1 min
user-configurable custom override as low as 30 s

At a 30 s interval against a 90 s timeout, up to 3 probes for the same provider run concurrently. For Antigravity each probe spawns its own copy of the 430 MB one-file agy_acp_server.exe, so three simultaneous PyInstaller extractions compete for disk and CPU. That is ~2,880 spawn attempts/day for a single provider.

This is distinct from #9650. #9650 is about the orphaned _MEI* extraction directories left behind when a probe is force-terminated. This issue is about concurrent overlapping probes, which remains a problem even with cleanup fixed correctly:

  • Redundant work — at most one probe result per interval can be useful.
  • Each Antigravity probe is a 430 MB unpack, so overlap multiplies sustained disk I/O and SSD wear.
  • Concurrent probes make the timeout self-fulfilling: three simultaneous extractions contend, so each is slower and more likely to exceed 90 s, which triggers another round.

It also interacts badly with #7230 — a timed-out probe cached as an error keeps the provider from ever reaching ready, so the loop never backs off.

Impact

Minor degradation

Version or commit

T3 Code (Alpha) 0.0.40.0, Windows x64

Environment

Windows 11 Pro 26200, T3 Code desktop 0.0.40.0, managed Antigravity 1.1.1 (agy_acp_server.exe, 430,801,616 bytes)

Logs or stack traces

# apps/server — AntigravityProvider
const checkProvider = fn$1("checkAntigravityProvider")(function* () {
  if (!settings.enabled) return yield* getSnapshot;
  const before = yield* get$3(metadata);
  const result = yield* options.probe.pipe(timeoutOption(HEALTH_CHECK_TIMEOUT), result$1);
  ...

const HEALTH_CHECK_TIMEOUT = "90 seconds";

# packages/shared/src/backgroundActivitySettings.ts
const DEFAULT_PROVIDER_HEALTH_REFRESH_INTERVAL = minutes(5);
PRESET_SETTINGS.performance.providerHealthRefreshInterval    = minutes(1);
PRESET_SETTINGS["battery-saver"].providerHealthRefreshInterval = minutes(15);
# overrides.providerHealthRefreshInterval is accepted verbatim, with no lower bound

# A single trace on this machine, one provider:
#   checkAntigravityProvider      x24
#   makeAntigravityAcpRuntime     x24
#   RpcClient.initialize          x24
#   effect-acp/AcpClient.make     x24
#   antigravityAuthSupport.handleStderr  x82

Suggested fix: clamp the effective interval to max(providerHealthRefreshInterval, HEALTH_CHECK_TIMEOUT) for process-based probes, or gate probes behind a per-provider semaphore so a scheduled tick is skipped while one is already in flight.

Screenshots, recordings, or supporting files

No response


Generated by Opus 5 xhigh

Activity

  1. juliusmarminge commented on Sep 16, 2026

    @juliusmarminge
    Member

    Triage

    Confirmed on current main (935c55b37) and on the reporter’s v0.0.40 line. This is a real config/perf bug, but the “up to 3 concurrent Effect probes” model is wrong. makeManagedServerProvider already serializes checkProvider per instance.

    The part that stands: providerHealthRefreshInterval and HEALTH_CHECK_TIMEOUT are independent, the UI will save 30s (or 0), and a short interval makes Antigravity’s 430 MB one-file spawn run far too often. That is enough to keep this open. It is not a wash of #9650 or #7230.

    What happens

    Antigravity’s health check is a full ACP spawn + initialize, capped at 90s (HEALTH_CHECK_TIMEOUT in apps/server/src/provider/Layers/AntigravityProvider.ts). The probe builds a disposable runtime (t3-code-provider-probe) and calls runtime.initialize(). On Windows that is the managed one-file binary (agy_acp_server.exe, ~430 MB).

    The refresh interval is a separate setting. Presets and the custom override have no relationship to that 90s cap:

    Setting Value
    HEALTH_CHECK_TIMEOUT 90s
    battery-saver 15 min
    balanced / DEFAULT_PROVIDER_HEALTH_REFRESH_INTERVAL 5 min
    performance 1 min (already < 90s)
    custom UI 0 or 30s+ (min={0}, step 30)

    resolveBackgroundActivitySettings copies overrides.providerHealthRefreshInterval verbatim. The schema is DurationFromMillis with no minimum. Background Activity and Providers → Advanced both use NumberField min={0}.

    So a custom 30s interval is legal, and the reporter does not even need it: performance is already 60s vs a 90s timeout.

    Concurrency — the report overstates this

    makeManagedServerProvider has had a 1-permit refreshSemaphore since the helper landed. Interval refresh, settings-driven refresh, the initial probe, and provider.refresh all go through applySnapshot with refreshSemaphore.withPermits(1).

    The periodic loop is also sequential: sleep the interval, then await refreshSnapshot(), then loop. The interval is the gap after a completed probe, not a fire-and-forget tick. A 30s interval against a 90s timeout therefore does not start three overlapping checkAntigravityProvider effects.

    What you actually get:

    • Fast probe (~seconds): one spawn every 30s (~2,880/day), sequential.
    • Probe that hits the 90s timeout: ~90s of work, then 30s sleep, then the next spawn (~720/day), still one Effect at a time.
    • Startup / settings / manual refresh while a probe is in flight: they queue on the semaphore, then run back-to-back. That is a burst of sequential 430 MB unpacks, not three parallel ones.
    • OS-level leftovers if timeoutOption interrupt + child.kill({ forceKillAfter: "1 second" }) returns before PyInstaller is actually gone: that is [Bug]: Antigravity health checks leave large _MEI folders in Windows temp #9650, not a missing scheduler lock.

    The checkAntigravityProvider x24 / makeAntigravityAcpRuntime x24 trace is consistent with sequential 30s ticks. handleStderr x82 is multiple stderr chunks per process, not 3 live runtimes.

    Related, not duplicates

    • #9650 — orphaned _MEI* dirs after Windows taskkill /T /F. Same spawn, different bug (cleanup). A short interval makes it worse; fixing cleanup does not add a floor, and a floor does not remove the force-kill leak.
    • #7230 / PR #7232 — a timed-out probe is published as status: "error" and cached. Distinct. fix(server): a provider probe timeout no longer marks the provider broken #7232 does not clamp the interval. A 30s loop plus sticky errors is why the reporter said these interact.

    No open PR clamps the interval or coalesces queued refreshes. Open PR #11657 (Fixes #9650) mitigates Antigravity spawn cost / orphans but does not address the interval floor.

    Suggested fix

    Keep this as a bugfix, not a redesign of provider health:

    1. Floor the effective interval for process-based probes: max(providerHealthRefreshInterval, HEALTH_CHECK_TIMEOUT) (or a documented minimum such as 90s). Apply it in makeManagedServerProvider / resolveServerBackgroundActivitySettings, not only in the NumberField, so API/legacy patches cannot sneak in 5s.
    2. Raise the UI minimum above 0 for “on” (keep 0 = disable). Step 30 with min={0} is how 30s gets saved today.
    3. Optional but useful: coalesce instead of queue. If a probe is in flight, skip the scheduled tick (withPermitsIfAvailable) so a 90s check is not followed by every refresh that piled up behind the semaphore.
    4. Tests: custom 30s + a 90s in-flight probe must not start a second checkProvider; a 5s persisted override must not survive resolve as 5s.

    Do not treat #7232 or a #9650 cleanup PR as closing this.

    Workaround: set Provider health interval to 0 (disables periodic probes), or use balanced / battery-saver. Disabling the Antigravity instance also stops the spawn loop.

    Classification: bug · accepted · minor–medium (config/perf; Windows Antigravity worst case)
    Labels: add bug, accepted, via-triage
    Discord tags: providers, performance, windows

  2. added
    bugSomething is broken or behaving incorrectly.
    acceptedfeature request accepted
    via-triageFiled through npx t3 triage
    on Sep 16, 2026
  3. fixfon commented on Oct 11, 2026

    @fixfon
    Contributor

    I'd like to take this. I'll follow the direction in the maintainer triage above and open a focused PR with before/after evidence, noting the model and harness used.

  4. fixfon commented on Oct 11, 2026

    @fixfon
    Contributor

    PR: #18210

    It floors the resolved interval at 30s (0 still disables) and normalizes the two Settings fields on commit. I used 30s rather than the 90s HEALTH_CHECK_TIMEOUT because #12008 removed the Antigravity spawn from the periodic probe after this was triaged, and a 90s floor would also change the performance preset. Presets are unchanged, and the in-flight tick skip (item 3) is left out. Details are in the PR description.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions