Skip to content

The Viewer's Overview shares one fleet-health read across its server cards, and the status bar stops re-measuring the store on every refresh (#4477) - #4482

Merged
erikdarlingdata merged 3 commits into
devfrom
fix/4477-viewer-store-cost
Sep 27, 2026
Merged

erikdarlingdata merged 3 commits into
devfrom
fix/4477-viewer-store-cost

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Refs #4477.

Why

A Darling Viewer session on a production store (43 servers), 4.5 minutes across the Overview, the fleet tabs and one server's tabs, measured from pg_stat_statements for the Viewer's role:

  • the Overview's fleet collection-health read (WITH parts AS …) ran 40 times. GetFleetCollectionHealthByServerAsync memoizes its result, but every server card loading at once raced the cold memo, and each started its own fleet-wide statement;
  • the status bar's SELECT pg_database_size(current_database()) ran 9 times, once per refresh, and it walks every file of the store.

The session's timings may be inflated by other load on that host at the time, and will be re-measured. The call counts don't depend on load.

What changes

PerformanceMonitor.Common.SingleFlightTtlCache<T> (new): a small TTL-memoized, single-flight gate for one async fetch.

  • Concurrent callers on a cold or expired cache share ONE fetch.
  • The fetch starts via Task.Run, so it can't complete inline before the shared in-flight field is assigned. A delegate that throws synchronously can't leave a finished, faulted task parked in the field.
  • The field is cleared, and the cache written, only by a lock-gated continuation that checks the task's own reference.
  • A caller's token cancels only that caller's WaitAsync on the shared task. The fetch itself runs with CancellationToken.None and its own command timeout, so one caller giving up never cancels the read for the others.

Its users:

  • ViewerDataService.GetFleetCollectionHealthByServerAsync reads through it, with the existing memo lifetime. N cards → one statement per refresh.
  • ViewerDataService.GetStoreSizeBytesAsync reads through it with a 5-minute lifetime (StoreSizeCacheLifetime). That's a coarse "about how big is the store" field, never a threshold. A null reading isn't cached, so a transient failure retries on the next call.

What's next

  • The two slowest Viewer reads in the same session are NOT changed here:
    • the Job History read (ViewerDataService.BuildJobHistorySql, whose WITH svr AS (SELECT DISTINCT ON (server_id) … FROM server_properties …) CTE is unbounded over all of server_properties; mean 8.5 s, max 13.8 s);
    • the collection-health summary reads (a max of 16–18 s, past the Viewer's 15 s client deadline).
  • Their plans are being captured on a production store first, and they get their own PR.

Test plan

  • Build: dotnet build Darling/Darling.Tests/Darling.Tests.csproj -c Release -p:EnableWindowsTargeting=true — 0 warnings, 0 errors.
  • Darling.Tests.SingleFlightTtlCacheTests (new, pure, no live Postgres): 6/6 GREEN, including a test that
    reproduces the wedge an inline start would cause (a synchronously-faulting fetch).
  • Darling.Tests.OverviewFleetHealthSingleFlightLiveTests and Darling.Tests.StoreSizeCacheLiveTests: GREEN
    against a live TimescaleDB rig with pg_stat_statements (the shared read still returns the seeded server's rows: 1 run, 1 success); RED against dev (12 calls / 2 calls where 1 is expected); RED against two mutations of the shared helper (disabling the in-flight join;
    removing the WaitAsync cancellation wrapper).
  • Darling.Tests.DocCommentHygieneTests: 77/77 GREEN.
  • Full touched-class run: 85/85 GREEN.

CHANGELOG

SECTION: Changed
ENTRY:

…ar's store size (#4477)

- ViewerDataService.GetFleetCollectionHealthByServerAsync gates a cold
  cache behind one in-flight fetch, so an Overview refresh's per-card
  lanes racing a cold 20s memo now share ONE fleet-wide statement
  instead of one per card.
- ViewerDataService.GetStoreSizeBytesAsync caches pg_database_size for
  5 minutes so the status bar's per-refresh walk of the whole store
  directory runs at most once per window.
- Added OverviewFleetHealthSingleFlightLiveTests and
  StoreSizeCacheLiveTests (statement-count + equivalence pins, not yet
  verified live against a rig — see the PR body's Test plan).
…completion (#4477)

The previous fix for the Overview's fleet-health read and the store-size
cache had two defects, found in review before either shipped:

- the fleet-health single-flight ran its fetch inline under the lock, so
  a fetch that completed (or faulted) before its first real await could
  race the assignment and wedge every later caller onto a stale
  completed task until the memo's TTL expired;
- neither cache let a waiter cancel its own wait without cancelling the
  shared fetch for every other caller racing it.

Both reads now go through one small SingleFlightTtlCache<T> helper
(PerformanceMonitor.Common): every fetch starts via Task.Run so it
cannot finish before the in-flight task is assigned, the field is
cleared by a lock-gated continuation that can only run after that
assignment, and a caller's own CancellationToken only cancels its own
WaitAsync on the shared task, never the fetch itself.

Pins: SingleFlightTtlCacheTests (6 pure tests: concurrent share-one-fetch,
a reproduction of the old inline pattern's wedge, synchronous and async
fault recovery, TTL expiry, and cancel-one-waiter). The two live tests
(OverviewFleetHealthSingleFlightLiveTests, StoreSizeCacheLiveTests) now
run GREEN against a real Postgres/TimescaleDB rig, RED against the
pre-#4477 commit (12 calls / 2 calls instead of 1), and RED against a
mutation of the shared helper (both the join-in-flight branch and the
WaitAsync call).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant