Repository navigation
The Viewer's Overview shares one fleet-health read across its server cards, and the status bar stops re-measuring the store on every refresh (#4477) - #4482
Merged
Conversation
…ar's store size (#4477) - ViewerDataService.GetFleetCollectionHealthByServerAsync gates a cold cache behind one in-flight fetch, so an Overview refresh's per-card lanes racing a cold 20s memo now share ONE fleet-wide statement instead of one per card. - ViewerDataService.GetStoreSizeBytesAsync caches pg_database_size for 5 minutes so the status bar's per-refresh walk of the whole store directory runs at most once per window. - Added OverviewFleetHealthSingleFlightLiveTests and StoreSizeCacheLiveTests (statement-count + equivalence pins, not yet verified live against a rig — see the PR body's Test plan).
…completion (#4477) The previous fix for the Overview's fleet-health read and the store-size cache had two defects, found in review before either shipped: - the fleet-health single-flight ran its fetch inline under the lock, so a fetch that completed (or faulted) before its first real await could race the assignment and wedge every later caller onto a stale completed task until the memo's TTL expired; - neither cache let a waiter cancel its own wait without cancelling the shared fetch for every other caller racing it. Both reads now go through one small SingleFlightTtlCache<T> helper (PerformanceMonitor.Common): every fetch starts via Task.Run so it cannot finish before the in-flight task is assigned, the field is cleared by a lock-gated continuation that can only run after that assignment, and a caller's own CancellationToken only cancels its own WaitAsync on the shared task, never the fetch itself. Pins: SingleFlightTtlCacheTests (6 pure tests: concurrent share-one-fetch, a reproduction of the old inline pattern's wedge, synchronous and async fault recovery, TTL expiry, and cancel-one-waiter). The two live tests (OverviewFleetHealthSingleFlightLiveTests, StoreSizeCacheLiveTests) now run GREEN against a real Postgres/TimescaleDB rig, RED against the pre-#4477 commit (12 calls / 2 calls instead of 1), and RED against a mutation of the shared helper (both the join-in-flight branch and the WaitAsync call).
erikdarlingdata
marked this pull request as ready for review
September 27, 2026 17:32
This was referenced Sep 27, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #4477.
Why
A Darling Viewer session on a production store (43 servers), 4.5 minutes across the Overview, the fleet tabs and one server's tabs, measured from
pg_stat_statementsfor the Viewer's role:WITH parts AS …) ran 40 times.GetFleetCollectionHealthByServerAsyncmemoizes its result, but every server card loading at once raced the cold memo, and each started its own fleet-wide statement;SELECT pg_database_size(current_database())ran 9 times, once per refresh, and it walks every file of the store.The session's timings may be inflated by other load on that host at the time, and will be re-measured. The call counts don't depend on load.
What changes
PerformanceMonitor.Common.SingleFlightTtlCache<T>(new): a small TTL-memoized, single-flight gate for one async fetch.Task.Run, so it can't complete inline before the shared in-flight field is assigned. A delegate that throws synchronously can't leave a finished, faulted task parked in the field.WaitAsyncon the shared task. The fetch itself runs withCancellationToken.Noneand its own command timeout, so one caller giving up never cancels the read for the others.Its users:
ViewerDataService.GetFleetCollectionHealthByServerAsyncreads through it, with the existing memo lifetime. N cards → one statement per refresh.ViewerDataService.GetStoreSizeBytesAsyncreads through it with a 5-minute lifetime (StoreSizeCacheLifetime). That's a coarse "about how big is the store" field, never a threshold. A null reading isn't cached, so a transient failure retries on the next call.What's next
ViewerDataService.BuildJobHistorySql, whoseWITH svr AS (SELECT DISTINCT ON (server_id) … FROM server_properties …)CTE is unbounded over all ofserver_properties; mean 8.5 s, max 13.8 s);Test plan
dotnet build Darling/Darling.Tests/Darling.Tests.csproj -c Release -p:EnableWindowsTargeting=true— 0 warnings, 0 errors.Darling.Tests.SingleFlightTtlCacheTests(new, pure, no live Postgres): 6/6 GREEN, including a test thatreproduces the wedge an inline start would cause (a synchronously-faulting fetch).
Darling.Tests.OverviewFleetHealthSingleFlightLiveTestsandDarling.Tests.StoreSizeCacheLiveTests: GREENagainst a live TimescaleDB rig with
pg_stat_statements(the shared read still returns the seeded server's rows: 1 run, 1 success); RED againstdev(12 calls / 2 calls where 1 is expected); RED against two mutations of the shared helper (disabling the in-flight join;removing the
WaitAsynccancellation wrapper).Darling.Tests.DocCommentHygieneTests: 77/77 GREEN.CHANGELOG
SECTION: Changed
ENTRY:
REF:
[The Viewer's Overview shares one fleet-health read across its server cards, and the status bar stops re-measuring the store on every refresh (#4477) #4482]: The Viewer's Overview shares one fleet-health read across its server cards, and the status bar stops re-measuring the store on every refresh (#4477) #4482