Skip to content

Daily Summary calendar caches closed days for an hour (#4232) - #4307

Merged
erikdarlingdata merged 6 commits into
devfrom
fix/4232-daily-summary-closed-day-cache
Sep 25, 2026
Merged

erikdarlingdata merged 6 commits into
devfrom
fix/4232-daily-summary-closed-day-cache

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Closes #4232.

Why

The Daily Summary calendar (WPF, every 1-minute tab refresh) and get_daily_summary_range (web Overview tab
through /api/read, MCP clients) recompute the whole displayed range on every poll, even though
every day but today (and briefly yesterday) is closed and its rows do not change. The issue measured this at
4.4 s cold / 0.41 s warm per poll for a 25-day range on a production store. Lite's own month read measured
94.6 ms warm on the same shape of query, already under the bar the issue set, so Lite needed no change
(issuecomment-5834845077) -- this PR is the whole fix for #4232.

What changes

  • New DailySummaryRangeCache<TRow> in PerformanceMonitor.Darling.Storage, next to DailySummarySql. It
    holds the CLOSED portion of one range read as one block for an hour. A refresh inside that hour re-runs the
    statement only over the days still open (today, plus yesterday during a two-hour post-midnight grace for a
    late collector run) and joins them to the block. After an hour, one refresh recomputes the whole range and
    repopulates it. The caller resolves the tier-routed statement text ONCE for the range as a whole and the
    cache reuses that same text for every sub-range it asks for, including the open-only re-read, folded into the
    cache key -- this stops a sub-range from routing to a different retention tier than the range it is part of.
  • The cache is bounded (d012ac8c). On every miss, it first drops every block older than BlockTtl (one
    hour), then evicts the oldest blocks until fewer than MaxBlocks (1,024) remain, before it adds the new
    block. Both removals use the TryRemove(KeyValuePair) overload, so a block that another caller just
    replaced is never removed. Before this, the key carried the range, so every distinct server/range/SQL
    combination added a block that nothing removed.
  • DarlingHealthReader.GetDailySummaryRangeAsync (service) routes through one static
    DarlingHealthReader.RangeCache, shared by the web viewer and MCP clients. ViewerDataService.ReadDailySummaryRangeAsync
    (WPF) keeps its own per-instance cache. Both readers cache the RAW per-day row from before the
    RateTiers/ReferenceUtc/DataState/RetentionHorizon stamp and re-apply that stamp fresh to every row,
    cached or fresh, on every call -- a deadlock-rate-threshold setting or the purge horizon can move inside the
    cache's one-hour lifetime.
  • The contract, plainly stated (The Daily Summary calendar recomputes a month of closed days from raw on every refresh: 0.55 s warm, 4.5 s cold per poll for 25 days that no longer change (WPF Daily Summary tab every minute, web server Overview every 60 s) #4232 ruling): a closed day's row can now be up to one hour stale by
    design. A row that lands late for a closed day -- an outage catch-up, a backfill -- shows within the hour,
    not immediately, where it showed immediately before this change. An explicit as_of timestamp always reads
    live and is never served from the cache.
  • Added DailySummaryRangeCache<TRow>.Clear() (public -- Storage's InternalsVisibleTo reaches
    Darling.Tests only, not the service) and DarlingHealthReader.ResetRangeCacheForTests() (internal, visible
    to Darling.Tests through the service's own InternalsVisibleTo), a test-only seam so a live test can force
    its next read to hit the store instead of a block an earlier call in the same process warmed.
  • Fixed a pre-existing doc-comment defect the full suite caught while running: GetDailySummaryRangeAsync's
    summary and RangeCache's summary sat back to back with no blank line between them, so both attached to
    RangeCache as one member's trivia (DocCommentHygieneTests.NoMemberCarriesTwoStackedSummaryBlocks).
    Reordered so each member carries exactly one summary block.

Census (#4232 item 2): every test reaching the cached range readers

Grepped every test file reaching GetDailySummaryRange, GetDailySummaryRangeAsync, or
ReadDailySummaryRangeAsync across Darling.Tests (the pattern also catches the Async spelling as a
substring): DailySummaryNotCarriedTests.cs, DailySummaryReadShapeTests.cs, DarlingDailySummaryRangeTests.cs,
McpFilterSemanticsLivePostgresTests.cs, McpToolGuideHeads.SqlCore.cs (a doc-comment mention, not a call),
ViewerCalendarRetentionPortTests.cs. Also checked every single-day GetDailySummary/GetDailySummaryAsync
call site, since that path funnels through the same cached range reader over a one-day window
(DarlingHealthReader.GetDailySummaryAsync calls GetDailySummaryRangeAsync for [day, day+1)).

Found and fixed three call sites that mutate a closed day's rows and then read the same server/range again
through the same store object -- each now calls ResetRangeCacheForTests() between the mutation and the
re-read, with a comment naming the one-hour contract and #4232:

  • McpFilterSemanticsLivePostgresTests.DailySummary_StopsPaintingPurgedDaysGreen_AndPublishesTheHorizon:
    inserts a deadlock on a 45-day-old closed day, then re-reads the same 60-day range -- this is the failure
    the brief named. The same test later DELETEs that row and reads the range a third time; same pattern, same
    fix, even though the assertions on that third read do not happen to check the specific value the delete
    changed.
  • DailySummaryNotCarriedTests.ASkippedDayBelowTheCeiling_ReadsNullAndIsNamed_AnEmptyDayIsAbsent_AndTheRepairClosesIt_AgainstDevPostgres:
    the materialization repair does not touch the old server's rows (the test's own comment says so), but the
    first read of that exact range had already warmed the block, so the closing assertion was being served from
    cache and never reaching the store at all. The reset restores it as a real check.

Checked and found no fix needed:

  • DarlingDailySummaryRangeTests.GetDailySummaryRangeAsync_ClosedDayCache_SecondReadDoesNotSeeANewRowOnAClosedDay_ButDoesOnToday
    -- this IS the cache's own pin, deliberately exercising the one-hour behavior.
  • DarlingDailySummaryRangeTests.TheCalendar_BandsEachDaySeparately_ShowsCollectionGaps_AndAnchors -- its
    anchored re-reads pass an explicit as_of, which bypasses the cache entirely (ruling item 5); its
    never-collected check already has its own server ID (this PR's earlier commit).
  • DailySummaryReadShapeTests.EveryDailyRead_ProbesCoverageOnce_HoweverManyRace_AgainstDevPostgres -- races
    reads, no mutation between them.
  • ViewerCalendarRetentionLivePostgresTests.Calendar_BandsAPurgedDayNoData_AndJudgesAgainstTheMcpHorizon_AgainstDevPostgres
    -- the two DarlingHealthReader reads use different fromDate values, so they are different cache keys, and
    the viewer's own per-instance cache reads the seeded rows on its first (and only) call.
    DarlingMcpHealthToolsTests's GetDailySummary call sites -- one read per scenario, no mutate-then-reread.

EXPLAIN (ANALYZE, BUFFERS): old path vs. new path (#4232 ruling item 9)

Seeded one server on the rig's probe database (schema copied from the migrated darlingtest, same
TimescaleDB/UTC rig CI uses) with 31 closed days of wait_stats/collection_log at a 5-minute collection
cadence, 15 positive-delta wait types per run (133,920 wait_stats rows, 8,928 collection_log rows), plus a
partial open "today" at the same cadence (2,880 more wait_stats rows). Ran DailySummarySql.RangeSql (the
raw, untiered statement -- the same text GetWindowSignalsAsync's single-day read already runs, and a fresh
rig with no continuous aggregates routes here anyway) through PREPARE/EXPLAIN (ANALYZE, BUFFERS) EXECUTE,
twice each, warm numbers below:

Path Window Rows out Execution time Planning time Buffers (shared hit) wait_stats scan
Old (every refresh recomputes the whole range) 31 days 32 139.14 ms 6.50 ms 2,050 Seq Scan, 136,800 rows
New (cache hit, open-only re-read) 1 day (today) 1 3.03 ms 4.16 ms 67 Bitmap Index Scan on idx_wait_stats_time, 2,880 rows

A roughly 46x drop in execution time and 31x drop in buffer hits at this seed, and the planner switches from a
full sequential scan of every row this server has ever reported to an index-bounded scan of the one open day.
The reduction is structural, not tied to this seed's size: the new path's cost is bounded by one day's rows
regardless of how wide the displayed range is or how much history the server carries, so a busier server or a
wider range only widens the gap.

Test plan

  • Darling.Tests.DailySummaryRangeCacheTests (9 tests, no database), unchanged from this PR's earlier
    commit: all pass.
  • dotnet build on Storage, Service, Viewer and Darling.Tests: 0 warnings, 0 errors.
  • Two new tests for the bound (d012ac8c), both shown to fail with the eviction code removed:
    GetRangeAsync_ExpiredBlock_EvictedOnNextMissForAnotherKey_AndOldKeyMissesAgain and
    GetRangeAsync_MoreThanMaxBlocksDistinctKeys_CountStaysCapped_OldestKeysDropped (1,034 distinct keys,
    the count stays at 1,024, and the oldest key misses again). DailySummaryRangeCacheTests 11/11,
    DocCommentHygieneTests 77/77 pass.
  • McpFilterSemanticsLivePostgresTests (4 tests, including the fixed
    DailySummary_StopsPaintingPurgedDaysGreen_AndPublishesTheHorizon): all pass against the rig.
  • DailySummaryNotCarriedTests, DarlingDailySummaryRangeTests, DailySummaryRangeCacheTests,
    ViewerCalendarRetentionPortTests, ViewerCalendarRetentionLivePostgresTests, DailySummaryReadShapeTests
    (26 tests): all pass against the rig.
  • git merge origin/dev: clean, no conflicts.
  • Full Darling.Tests.exe suite, once, against a freshly created darlingtest on the UTC rig:
    Total: 14059, Errors: 0, Failed: 3 (at first pass), Skipped: 49 (environment-gated, e.g.
    DARLING_TEST_PGRUNTIME/jsonlog targets this lane did not set), Not Run: 1, Time: 960s.
    All 3 failures
    investigated:
    - DocCommentHygieneTests.NoMemberCarriesTwoStackedSummaryBlocks -- real, fixed above (own commit).
    - ServerListAndSummaryPlanShapeTests.TheShippedReads_TouchFarFewerChunks_AndReturnTheSameNewestCollection_AgainstDevPostgres
    and CaptureDownChunkOrderTests.TheShippedRead_ExecutesOnlyTheNewestChunk_AndTheNewestRunDecides_AgainstDevPostgres
    -- both plan-shape/chunk-order tests unrelated to the daily-summary files this PR touches. Re-ran both,
    plus the doc-comment test, alone on a freshly recreated darlingtest: all 81 tests in those three
    classes passed (Total: 81, Failed: 0). Not reproducible in isolation, so the cause was cross-test
    state from the single big run, not this PR's code; not filed.
  • EXPLAIN (ANALYZE, BUFFERS) table above (ruling item 9).

What the coordinator should double-check

  • The two chunk-order test failures that only reproduced inside the full run, never alone -- worth a look if
    they recur in CI, since a flake that only shows up under the full suite's ordering/load is exactly the kind
    the wave-boundary census watches for.

CHANGELOG entry

SECTION: Fixed
ENTRY:

🤖 Generated with Claude Code

https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ

erikdarlingdata and others added 6 commits September 25, 2026 11:29
DailySummaryRangeCache<TRow> holds the closed-day portion of one
daily-summary range read as one block for an hour; a refresh inside
that hour re-runs the statement only over the still-open days (today,
and yesterday during a two-hour post-midnight grace) and joins them to
the block. Lives in Storage next to DailySummarySql so both the WPF
viewer and the service can share it without seeing each other.

Checked ruling item 7 first: no column DailySummarySql.RangeSqlFor
returns depends on more than its own day (every CTE groups strictly
within its own WHERE-bounded rows; no window function spans days), so
there is no band or rank to protect from the cache. TRow is still the
row from BEFORE any later banding/threshold/horizon stamp, so callers
re-apply that judgment fresh after every read, cached or not.

Unit pins (no database): cached-vs-fresh equality across a late row
landing for yesterday inside the grace window, a second refresh
reading only the open sub-range through the runRange seam, the
one-hour block TTL and two-hour grace with a moved clock, an explicit
end time skipping the cache every time, and a dedicated pin for the
case where the grace boundary crosses mid-TTL (proved against the
unguarded shape: without the closed-end boundary check, yesterday's
row is silently dropped from the join).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
DarlingHealthReader.GetDailySummaryRangeAsync (get_daily_summary_range,
shared by the web /api/read mirror and MCP clients) and
ViewerDataService.ReadDailySummaryRangeAsync (the WPF Daily Summary
calendar) now route their range reads through DailySummaryRangeCache
instead of reading every day fresh on every refresh.

Both resolve the tier-routed SQL text ONCE per call and reuse that
same text for every sub-range the cache asks for (including the
open-days-only re-read), so a 30-day range that routes to the hourly
rollup cannot have its "today" slice quietly re-resolve to raw. Each
reader splits its row read into a raw parse (what gets cached) and a
banding/retention stamp (RateTiers/ReferenceUtc/DataState/
RetentionHorizon for the service, DataState/HealthBand for the
viewer) applied fresh to every row, cached or not, on every call.

The service keeps ONE static cache (DarlingHealthReader.RangeCache),
keyed in part by the NpgsqlDataSource reference so distinct stores
never collide. The viewer keeps a per-instance cache field, matching
the existing GetRollupAvailabilityAsync pattern. get_daily_summary_range
now passes asOfNow: as_of is null, so an explicit as_of always reads
live per ruling item 5.

Also: a pre-existing live test (DarlingDailySummaryRangeTests) shared
one server ID between an empty-store check and a seeded-range check;
the empty check now warms the cache's block for that server/range
first, so the seeded read (same block, same hour) saw the stale empty
block instead of the just-seeded rows. Split it onto its own server ID
-- the "nothing ever collected" case does not need to share
infrastructure with the gap-and-band case, and doing so was papering
over two different stores as one. Added a new live pin that seeds a
closed day and an open day, reads, seeds a second row on each, reads
again, and asserts the closed day is unchanged (served from the cached
block) while today picks up its new row (never cached).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
McpFilterSemanticsLivePostgresTests.DailySummary_StopsPaintingPurgedDaysGreen
mutates a closed day (45 days old) between two range reads of the same
server/range through the same store object, then asserts the second read
reflects the mutation immediately. The #4232 cache serves that read from its
one-hour block instead, by design (a real backfill shows up within the hour,
not on the next statement) -- so the test was testing the cache's own
contract, not a product bug.

Add DailySummaryRangeCache<TRow>.Clear() (public: Storage's
InternalsVisibleTo reaches Darling.Tests only, not the service) and
DarlingHealthReader.ResetRangeCacheForTests() (internal, visible to
Darling.Tests through the service's own InternalsVisibleTo). Call it between
each closed-day mutation and the re-read that expects to see it, in every
test the census found doing this:

- McpFilterSemanticsLivePostgresTests: the deadlock insert before the
  `survived` read, and the deadlock delete before the `shortened` read
  (same test, two separate mutate/reread pairs on the same ghost day).
- DailySummaryNotCarriedTests.ASkippedDayBelowTheCeiling...: the repair
  pass doesn't touch the old server's rows, but the first read at the top
  of the test had already warmed that exact range's block, so the closing
  assertion was being served from cache and never reaching the store at
  all -- reset restores it as a real check.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…4232)

The full Darling.Tests.exe run flagged DocCommentHygieneTests.NoMember
CarriesTwoStackedSummaryBlocks: GetDailySummaryRangeAsync's own summary and
RangeCache's summary sat back to back with no blank line between them, so
both attached to RangeCache as trivia. Predates this branch's cache work
(present in a75cd1d, before this lane's changes); reordered so each member
carries exactly one summary block, keeping the new ResetRangeCacheForTests
doc in place.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
DailySummaryRangeCache never removed a block: the key carries the
range, so every days_back value, every server and every day added an
entry nothing replaced. Trim expired blocks (past BlockTtl) and cap
the total at MaxBlocks (1024), evicting the oldest first, on every
cache miss.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
@erikdarlingdata

Copy link
Copy Markdown
Owner Author

Pushed the bound fix to fix/4232-daily-summary-closed-day-cache (commit d012ac8, on top of 8ca50b9).

The change

Darling/PerformanceMonitor.Darling.Storage/DailySummaryRangeCache.cs:

  • GetRangeAsync's miss path (line ~161, just before the block is inserted) now trims before writing:
    • A foreach over _blocks removes every entry whose nowUtc - ComputedAtUtc >= BlockTtl, via the
      TryRemove(KeyValuePair) conditional overload, so a concurrent caller's fresher replacement is never
      clobbered.
    • A while (_blocks.Count >= MaxBlocks && !_blocks.IsEmpty) loop then evicts the single oldest block
      (MinBy(entry => entry.Value.ComputedAtUtc)) until the dictionary is back under the cap, breaking out if
      a concurrent TryRemove already lost the race (the next miss trims again).
  • public const int MaxBlocks = 1024; (line 69), with a doc comment: the fleet case is one or two ranges per
    server, and a block holds at most 366 rows, so this is generous headroom.
  • internal int Count => _blocks.Count; (line 75), for the tests.
  • The class doc comment (lines 24-27) now states blocks expire after BlockTtl and the cache holds at most
    MaxBlocks, evicting the oldest first.

Tests

Darling/Darling.Tests/DailySummaryRangeCacheTests.cs, both new, both revert-proven (temporarily replaced the
miss-path insert with the bare _blocks[key] = new CachedBlock(...) line, no trimming, rebuilt, and confirmed
each failed; then restored the fix and confirmed all pass again):

  • GetRangeAsync_ExpiredBlock_EvictedOnNextMissForAnotherKey_AndOldKeyMissesAgain: builds a block for server 1,
    moves the injected clock past BlockTtl, then misses on server 2 (same range, different key). Asserts
    Count == 1 after that second call (server 1's stale block swept, only server 2's remains) and that a
    follow-up read for server 1 calls runRange again over the whole range, not a cached hit. Without the fix,
    Count is 2 after the second call: Assert.Equal() Failure: Values differ.
  • GetRangeAsync_MoreThanMaxBlocksDistinctKeys_CountStaysCapped_OldestKeysDropped: fills MaxBlocks + 10
    distinct keys (1034), one second apart so only the count cap fires, not the TTL sweep. Asserts
    Count == MaxBlocks afterward, and that a re-read for server 1 (the oldest, evicted first) is a miss
    (runRange called once, over the whole range) while the cap still holds after the re-add. Without the fix,
    Count is 1034: Assert.Equal() Failure: Values differ.
  • The 9 pre-existing tests in the same file pass unchanged; their TTL/grace-boundary logic reuses the same key
    each call in every scenario that matters, so the new trim (which only affects OTHER keys, or fires on a miss
    that overwrites the exact key it inspects) doesn't change any of their outcomes.

Test totals (all against commit d012ac8, Debug build, 0 warnings / 0 errors)

  • DailySummaryRangeCacheTests: Total 11, Failed 0.
  • DarlingDailySummaryRangeTests: Total 2, Failed 0 (2 skipped -- these are live PostgreSQL tests; no rig was
    started for this change, unit tests only, per the brief).
  • DocCommentHygieneTests: Total 77, Failed 0 (covers the new/changed doc comments in
    DailySummaryRangeCache.cs -- summary balance and cref resolution both clean).

Full suite not run here; CI runs it. Skipped the plain-English pass on this comment per the brief.

🤖 Generated with Claude Code

https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant