Skip to content

The fleet overview reads each server's newest sample instead of every retained row, and the web viewer asks for it once per refresh (#3895) - #3947

Merged
erikdarlingdata merged 1 commit into
devfrom
feature/3895-fleet-overview-reads
Sep 23, 2026
Merged

erikdarlingdata merged 1 commit into
devfrom
feature/3895-fleet-overview-reads

Conversation

@erikdarlingdata

Copy link
Copy Markdown
Owner

Closes #3895.

Why

Every get_fleet_overview call and every /api/fleet request read whole tables to answer "what is each server's newest value":

  • Four newest-row reads were unbounded. FleetCpuSql, FleetMemorySql and FleetThreadsSql were DISTINCT ON (server_id) over the whole view, and FleetMemoryPressureSql was a MAX(collection_time) over the whole view joined back to it. Each one read, decompressed and sorted every retained row to keep one per server.
  • Two incident counts had no chunk bound. FleetBlockingSql and FleetDeadlockSql were bounded on the event's own timestamp. No chunk is partitioned on that column, so both opened every retained chunk to count one hour.
  • The last-collection read aggregated two days of the biggest table. FleetLastCollectionSql was bounded to 48 hours, but as a GROUP BY it still aggregated every collector run in that window (490,940 rows on DARLING01) to get about ten timestamps.
  • The web viewer asked twice. Every 60 s poll re-renders the sidebar and the current page in one pass, and both fetched /api/fleet. Every visible tab paid for the whole overview twice a minute.

The WPF viewer's Overview has its own copies of the same patterns. Its sidebar freshness read was a GROUP BY over the whole retained collection_log on every refresh tick, and its fleet totals and per-server windowed counts were also bounded on event time only. Lite's Overview card read its newest CPU, memory and last collection through views that union the hot table with every archived parquet month.

I confirmed the diagnosis on DARLING01 (9 servers, 18 daily chunks) and at production shape on my rig: 43 SQL Servers plus 5 PostgreSQL targets, a never-collected server and a dark one, 30 days, and 1.94 M rows per table with 28 compressed and 3 uncompressed chunks. There the old CPU read alone took 1,234 ms to execute.

What changes

Fleet reader (serves both /api/fleet and get_fleet_overview)

  • The four newest-row reads are now a per-server CROSS JOIN LATERAL (... WHERE server_id = s.server_id ORDER BY collection_time DESC [, sample_time DESC] LIMIT 1), driven from the enabled registry.
    • TimescaleDB's ordered ChunkAppend stops in the newest chunk for every server that is collecting. collection_time still leads the ordering, so the frame argument and LatestCpuReadShapeSqlTests hold.
    • Memory pressure finds the newest instant with that probe, then sums the pools at that instant by equality, which runtime chunk exclusion resolves to one chunk.
    • The driver skips servers whose registry row says they are PostgreSQL (SqlServerCollectedTargetSql). Such a target never writes these tables, and probing it walks every chunk to find nothing: 16 of the 17 ms the memory read cost on DARLING01. A NULL or unrecognised engine kind is still probed. The SQL normalises the kind the way MonitoredEngineKind.IsPostgres does, and a live test holds the two to the same answer for every token and spelling.
  • FleetBlockingSql and FleetDeadlockSql gain collection_time >= $3. $3 is EventWindowFloor.For(windowStart): the window start minus one chunk (the new Storage/EventWindowFloor.cs).
    • A row is collected after its event, apart from clock skew. Measured on DARLING01: blocked-process reports land 4.8 to 62 s after their event, a DMV snapshot's event time is its collection time, and the worst deadlock sat 27 ms after its collection. A one-day allowance keeps every such row.
    • There is deliberately no upper bound: a late catch-up collection is still an event in the window.
  • FleetLastCollectionSql is a per-server LIMIT 1 over the same 48 hours. The window is kept, so a server dark for longer falls out exactly as before.

WPF viewer

  • ServerFreshnessSql (sidebar dots and Manage Servers) becomes a per-server probe over every registry row, disabled servers included, unbounded as before.
  • FleetTotalsSql and the windowed halves of ServerSummaryBlockingSql and ServerSummaryDeadlockSql carry the same floor.
  • The two "newest event ever" reads keep no bound, because a floor would change their answer.

MCP get_server_summary

  • The per-server twin of the fleet card, DarlingHealthReader. Its windowed blocking and deadlock counts carry the same floor.

Web viewer

  • util.js adds apiGetFleet(). Callers that ask while a request is in flight share that one request, and nothing outlives the response, so no page renders an older roll-up than before.
  • Each caller parses the body itself, so no page can mutate another page's cards.
  • All seven /api/fleet callers use it: the sidebar, the fleet page, the server page, both composer loaders and both custom-view reads.

Lite (parity)

  • GetServerSummaryAsync reads its three newest rows from the hot table first and falls back to the v_ view only on a miss. ArchiveService archives only rows older than its cutoff, so a server's newest row is hot whenever it has any hot row.

The rows returned are unchanged. The one row the floor can refuse is an event collected more than a day before its own timestamp, meaning a monitored clock a day ahead of the collector's. The live test pins that as the documented boundary.

No migration, no index, no store-side setting.

Test plan

  • DARLING01, EXPLAIN (ANALYZE, BUFFERS) of the shipped SQL, old against new (warm, unprepared, second run of each; timings are plan + execution):
Read Before After Buffers
FleetCpuSql 2.3 + 460.3 ms 3.0 + 0.6 ms 2,906 → 28
FleetMemorySql 3.3 + 70.7 ms 3.2 + 0.7 ms 2,194 → 26
FleetMemoryPressureSql 8.0 + 256.6 ms 15.9 + 2.3 ms 9,537 → 44
FleetThreadsSql 3.3 + 59.7 ms 3.8 + 0.6 ms 2,506 → 28
FleetBlockingSql 10.6 + 113.0 ms 0.4 + 0.1 ms 88 → 12
FleetDeadlockSql 1.1 + 140.1 ms 0.2 + 0.04 ms 98 → 6
FleetLastCollectionSql 0.8 + 187.1 ms 0.8 + 0.3 ms 9,507 → 38
Fleet sub-reads, total 1,316.8 ms 32.0 ms 26,836 → 182
Viewer ServerFreshnessSql 4.8 + 297.7 ms 11.7 + 1.3 ms 9,751 → 184

An earlier cold-session run measured 40 to 63 ms of planning per unbounded read, and 2 to 3 ms for the probes.

  • Production-shape rig (43 SQL Servers, 30 days, 1.94 M rows per table):
    • CPU: 1,233.7 ms → 3.0 ms to execute, 2.0 ms to plan.
    • Memory probe: 2.6 ms.
    • Last collection: 160 ms → 1 ms.
  • Row equality on DARLING01: old and new rows are identical for CPU, memory, memory pressure, threads, last collection and the viewer's freshness read.
  • MCP before, on DARLING01's current build: get_fleet_overview(hours_back=1) took 2,712 ms on the first call (collection-health scan at age 0), then 531 to 779 ms warm. The after number needs the deploy.
  • Web, DOM-shim run of the shipped modules with a stubbed fetch:
    • /api/fleet requests per 60 s tick: fleet page 2 → 1, server page 2 → 1, saved view 2 → 1.
    • Initial load: 2 → 1.
    • apiGetFleet under node: concurrent callers join one request, each gets its own parsed copy, a later caller sends a fresh request, a transport failure comes back as an error to every joined caller and is not memoized, and results classify exactly like apiGet.
  • Lite, local store with June to September archives, per server:
    • CPU: 5.2-12.5 ms → 1.1-1.7 ms.
    • Memory: 4.2-7.4 ms → 0.8-1.5 ms.
    • Last collection: 16.6-27.6 ms → 0.5-0.8 ms.
  • New tests:
    • FleetReadsAreBoundedByTheFleetTests: every public fleet statement is bounded per server or by a partition window, with the four replaced shapes as positive controls. Every event-table scan carries the floor. The EventWindowFloor arithmetic is pinned. No JS file fetches /api/fleet except through apiGetFleet.
    • FleetReadsAreBoundedLivePostgresTests (live):
      • Old against new row equality over a seeded week with every sentinel shape: a tied CPU batch, pools summed at the newest instant, a five-day-dark server, a never-collected server, a disabled server, an unstamped engine, a PostgreSQL target with a planted row, and late-collected and clock-ahead incidents.
      • The overview seam, through the real reader.
      • With TimescaleDB, the plan: the probes return a handful of rows where DISTINCT ON streamed the week, and the floored count plans only the window's chunks.
      • The engine filter against the C# decoder.
    • Lite OverviewNewestReadsHotFirstTests: the archive-only fallback, and a hot server's card still comes back whole after its parquet archive is deleted, which the view-reading shape cannot survive.
  • Red-watch: reverting FleetMemorySql to DISTINCT ON, dropping the deadlock floor and restoring one direct fetch fails 8 tests.
  • Pins updated deliberately: DarlingFleetReaderSqlTests, DarlingMcpHealthToolsTests, ViewerServerChromeTests, ViewerW2aTests, ViewerFleetRollupTests, LatestCpuReadShapeSqlTests (doc, plus a positive control for the lateral shape) and Lite's TimeHonestyRungTests.
  • Full Darling.Tests with DARLING_TEST_PG: 12,470 tests and 1 failure, ForcePlanFailuresAccessPathTests.TheShippedRead_PlansAsAnIndexOnlyScanOnTheCoveringIndex. That failure predates this branch: a full run of dev at e507aa2d on the same rig fails it too, and it passes alone. Filed as ForcePlanFailuresAccessPathTests' index-only-scan assertion fails in a full live run, passes alone #3945.
  • Full Lite.Tests: 5,125 tests and 1 failure, AnalysisPassTokenThreadingTests.TheReadLockWaitIsAbandonableWhileAWriterHoldsIt. It belongs to the known AnalysisPass threading flake family and passes alone.
  • After deploy, re-time get_fleet_overview and /api/fleet on DARLING01.

Follow-ups filed

🤖 Generated with Claude Code

https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif

… retained row (#3895)

Every get_fleet_overview call and every /api/fleet request ran four
"latest value per server" reads as DISTINCT ON (or a MAX joined back)
over the whole table, and counted blocking and deadlocks bounded only
on the event's own timestamp, which no chunk is partitioned on. Both
grew with servers x retained days. The web viewer asked for the whole
overview twice per 60 s poll per tab.

- DarlingFleetReader: the CPU, memory, memory-pressure and threads
  reads become a per-server LATERAL ... LIMIT 1 driven from the enabled
  registry. PostgreSQL targets are skipped because they never write
  those tables. The incident counts carry a collection_time floor one
  chunk before the window (EventWindowFloor). The 48-hour last-collection
  read becomes a per-server probe that keeps its window.
- WPF viewer: the sidebar freshness read becomes a per-registry-server
  probe. The fleet totals and the per-server windowed counts carry the
  floor. get_server_summary's windowed counts carry it too.
- Web: apiGetFleet() shares one in-flight /api/fleet request among its
  callers, and each caller parses the body itself. A poll tick now sends
  one fleet request, not two.
- Lite: the Overview card's newest-row reads use the hot table first
  and fall back to the archive view.

DARLING01: fleet sub-reads 1,316.8 ms / 26,836 buffers -> 32.0 ms / 182,
with identical rows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif
@erikdarlingdata
erikdarlingdata enabled auto-merge (squash) September 23, 2026 00:17
@erikdarlingdata
erikdarlingdata merged commit 38b51d3 into dev Sep 23, 2026
15 of 16 checks passed
@erikdarlingdata
erikdarlingdata deleted the feature/3895-fleet-overview-reads branch September 23, 2026 00:22
erikdarlingdata added a commit that referenced this pull request Sep 23, 2026
…3920, #3927, #3931, #3932, #3940, #3942, #3946, #3947, #3950, #3952, #3955, #3956, #3957, #3964, #3965, #3966, #3968, #3972, #3975, #3979, #3980, #3981, #3983, #3984, #3985) (#3989)

The wave's fix PRs deliberately carried no CHANGELOG edits (parallel-agent hot-spot protocol); each agent reported its entry and this commit lands them together, byte-verified against origin/dev.


Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant