Repository navigation
The fleet overview reads each server's newest sample instead of every retained row, and the web viewer asks for it once per refresh (#3895) - #3947
Merged
Conversation
… retained row (#3895) Every get_fleet_overview call and every /api/fleet request ran four "latest value per server" reads as DISTINCT ON (or a MAX joined back) over the whole table, and counted blocking and deadlocks bounded only on the event's own timestamp, which no chunk is partitioned on. Both grew with servers x retained days. The web viewer asked for the whole overview twice per 60 s poll per tab. - DarlingFleetReader: the CPU, memory, memory-pressure and threads reads become a per-server LATERAL ... LIMIT 1 driven from the enabled registry. PostgreSQL targets are skipped because they never write those tables. The incident counts carry a collection_time floor one chunk before the window (EventWindowFloor). The 48-hour last-collection read becomes a per-server probe that keeps its window. - WPF viewer: the sidebar freshness read becomes a per-registry-server probe. The fleet totals and the per-server windowed counts carry the floor. get_server_summary's windowed counts carry it too. - Web: apiGetFleet() shares one in-flight /api/fleet request among its callers, and each caller parses the body itself. A poll tick now sends one fleet request, not two. - Lite: the Overview card's newest-row reads use the hot table first and fall back to the archive view. DARLING01: fleet sub-reads 1,316.8 ms / 26,836 buffers -> 32.0 ms / 182, with identical rows. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif
erikdarlingdata
enabled auto-merge (squash)
September 23, 2026 00:17
erikdarlingdata
added a commit
that referenced
this pull request
Sep 23, 2026
…3920, #3927, #3931, #3932, #3940, #3942, #3946, #3947, #3950, #3952, #3955, #3956, #3957, #3964, #3965, #3966, #3968, #3972, #3975, #3979, #3980, #3981, #3983, #3984, #3985) (#3989) The wave's fix PRs deliberately carried no CHANGELOG edits (parallel-agent hot-spot protocol); each agent reported its entry and this commit lands them together, byte-verified against origin/dev. Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This was referenced Sep 25, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #3895.
Why
Every
get_fleet_overviewcall and every/api/fleetrequest read whole tables to answer "what is each server's newest value":FleetCpuSql,FleetMemorySqlandFleetThreadsSqlwereDISTINCT ON (server_id)over the whole view, andFleetMemoryPressureSqlwas aMAX(collection_time)over the whole view joined back to it. Each one read, decompressed and sorted every retained row to keep one per server.FleetBlockingSqlandFleetDeadlockSqlwere bounded on the event's own timestamp. No chunk is partitioned on that column, so both opened every retained chunk to count one hour.FleetLastCollectionSqlwas bounded to 48 hours, but as aGROUP BYit still aggregated every collector run in that window (490,940 rows on DARLING01) to get about ten timestamps./api/fleet. Every visible tab paid for the whole overview twice a minute.The WPF viewer's Overview has its own copies of the same patterns. Its sidebar freshness read was a
GROUP BYover the whole retainedcollection_logon every refresh tick, and its fleet totals and per-server windowed counts were also bounded on event time only. Lite's Overview card read its newest CPU, memory and last collection through views that union the hot table with every archived parquet month.I confirmed the diagnosis on DARLING01 (9 servers, 18 daily chunks) and at production shape on my rig: 43 SQL Servers plus 5 PostgreSQL targets, a never-collected server and a dark one, 30 days, and 1.94 M rows per table with 28 compressed and 3 uncompressed chunks. There the old CPU read alone took 1,234 ms to execute.
What changes
Fleet reader (serves both
/api/fleetandget_fleet_overview)CROSS JOIN LATERAL (... WHERE server_id = s.server_id ORDER BY collection_time DESC [, sample_time DESC] LIMIT 1), driven from the enabled registry.collection_timestill leads the ordering, so the frame argument andLatestCpuReadShapeSqlTestshold.SqlServerCollectedTargetSql). Such a target never writes these tables, and probing it walks every chunk to find nothing: 16 of the 17 ms the memory read cost on DARLING01. A NULL or unrecognised engine kind is still probed. The SQL normalises the kind the wayMonitoredEngineKind.IsPostgresdoes, and a live test holds the two to the same answer for every token and spelling.FleetBlockingSqlandFleetDeadlockSqlgaincollection_time >= $3.$3isEventWindowFloor.For(windowStart): the window start minus one chunk (the newStorage/EventWindowFloor.cs).FleetLastCollectionSqlis a per-serverLIMIT 1over the same 48 hours. The window is kept, so a server dark for longer falls out exactly as before.WPF viewer
ServerFreshnessSql(sidebar dots and Manage Servers) becomes a per-server probe over every registry row, disabled servers included, unbounded as before.FleetTotalsSqland the windowed halves ofServerSummaryBlockingSqlandServerSummaryDeadlockSqlcarry the same floor.MCP
get_server_summaryDarlingHealthReader. Its windowed blocking and deadlock counts carry the same floor.Web viewer
util.jsaddsapiGetFleet(). Callers that ask while a request is in flight share that one request, and nothing outlives the response, so no page renders an older roll-up than before./api/fleetcallers use it: the sidebar, the fleet page, the server page, both composer loaders and both custom-view reads.Lite (parity)
GetServerSummaryAsyncreads its three newest rows from the hot table first and falls back to thev_view only on a miss.ArchiveServicearchives only rows older than its cutoff, so a server's newest row is hot whenever it has any hot row.The rows returned are unchanged. The one row the floor can refuse is an event collected more than a day before its own timestamp, meaning a monitored clock a day ahead of the collector's. The live test pins that as the documented boundary.
No migration, no index, no store-side setting.
Test plan
FleetCpuSqlFleetMemorySqlFleetMemoryPressureSqlFleetThreadsSqlFleetBlockingSqlFleetDeadlockSqlFleetLastCollectionSqlServerFreshnessSqlAn earlier cold-session run measured 40 to 63 ms of planning per unbounded read, and 2 to 3 ms for the probes.
get_fleet_overview(hours_back=1)took 2,712 ms on the first call (collection-health scan at age 0), then 531 to 779 ms warm. The after number needs the deploy./api/fleetrequests per 60 s tick: fleet page 2 → 1, server page 2 → 1, saved view 2 → 1.apiGetFleetunder node: concurrent callers join one request, each gets its own parsed copy, a later caller sends a fresh request, a transport failure comes back as an error to every joined caller and is not memoized, and results classify exactly likeapiGet.FleetReadsAreBoundedByTheFleetTests: every public fleet statement is bounded per server or by a partition window, with the four replaced shapes as positive controls. Every event-table scan carries the floor. TheEventWindowFloorarithmetic is pinned. No JS file fetches/api/fleetexcept throughapiGetFleet.FleetReadsAreBoundedLivePostgresTests(live):DISTINCT ONstreamed the week, and the floored count plans only the window's chunks.OverviewNewestReadsHotFirstTests: the archive-only fallback, and a hot server's card still comes back whole after its parquet archive is deleted, which the view-reading shape cannot survive.FleetMemorySqltoDISTINCT ON, dropping the deadlock floor and restoring one direct fetch fails 8 tests.DarlingFleetReaderSqlTests,DarlingMcpHealthToolsTests,ViewerServerChromeTests,ViewerW2aTests,ViewerFleetRollupTests,LatestCpuReadShapeSqlTests(doc, plus a positive control for the lateral shape) and Lite'sTimeHonestyRungTests.Darling.TestswithDARLING_TEST_PG: 12,470 tests and 1 failure,ForcePlanFailuresAccessPathTests.TheShippedRead_PlansAsAnIndexOnlyScanOnTheCoveringIndex. That failure predates this branch: a full run of dev ate507aa2don the same rig fails it too, and it passes alone. Filed as ForcePlanFailuresAccessPathTests' index-only-scan assertion fails in a full live run, passes alone #3945.Lite.Tests: 5,125 tests and 1 failure,AnalysisPassTokenThreadingTests.TheReadLockWaitIsAbandonableWhileAWriterHoldsIt. It belongs to the known AnalysisPass threading flake family and passes alone.get_fleet_overviewand/api/fleeton DARLING01.Follow-ups filed
cpu_scheduler_statsstores several different snapshots under onecollection_time(31 collections in 24 h on SQL2025, up to 13 rows each). Found while checking ties for the equality proof.🤖 Generated with Claude Code
https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif