Skip to content

The Daily Summary calendar recomputes a month of closed days from raw on every refresh: 0.55 s warm, 4.5 s cold per poll for 25 days that no longer change (WPF Daily Summary tab every minute, web server Overview every 60 s) #4232

Description

@erikdarlingdata

Problem

Both calendars rebuild a whole range of days on every refresh:

  • WPF: the Daily Summary tab re-reads the displayed month (GetDailySummaryRangeAsync(month start, month end)) on every server-tab auto-refresh (1 min).
  • Web: the server page's default Overview tab re-reads get_daily_summary_range with days_back: 30 on every 60 s poll.

The range statement is DailySummarySql. Its only tier-routed CTE is queries (#3905: "no other CTE has a tier to route to"). The other eight re-aggregate raw rows for every day in the range, every time:

  • wait_per_type: every wait type × every collection, the heaviest;
  • collection_log;
  • cpu_utilization_stats;
  • deadlocks, blocked_process_reports, dmv_blocking_snapshots;
  • memory_pressure_events;
  • config_alert_log.

Every day before today is closed. Its rows don't change except for a late collection landing in the first minutes after midnight, yet each poll recomputes all of them to redraw one open day.

Measured

Production SQL Server store B (42 servers), 2026-09-25, the busiest server by executions, the WPF month range (Sep 1 → Oct 1, 25 days of data), EXPLAIN (ANALYZE, BUFFERS), read-only. The queries CTE (which routes to the hourly rollup) was stubbed out, so this is exactly the non-routed part:

run exec planning buffers
cold 4,406 ms 144 ms 49.3 k (9.9 k read, 3.8 s I/O)
warm 410 ms 140 ms 49.9 k
wait_per_type CTE alone, warm 281 ms 3.5 ms 30.8 k

Earlier on the web viewer: get_daily_summary_range took 0.96 s on the server Overview tab, and the store's viewer-role statistics showed the wait_per_type statement at a mean of 5.9 s and a max of 11.3 s.

Where (origin/dev)

  • Darling/PerformanceMonitor.Darling.Storage/DailySummarySql.cs: RangeSql ~:33 (the per-source CTEs ~:34–135); RangeSqlFor ~:562–620, which routes only queries.
  • Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.DailySummary.cs: LoadDailySummaryAsync ~:33 → LoadCalendarMonthAsync ~:44–62, on the per-server auto-refresh.
  • Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.DailySummary.cs ~:71–125.
  • Web: Darling/PerformanceMonitor.Darling.Service/wwwroot/js/pages/server-tabs.js ~:523 (days_back: 30, Overview tab); MCP get_daily_summary_range (Mcp/DarlingMcpHealthTools.cs ~:183).

Fix shape

Recompute only the open days. Either of these works, and they combine:

  1. Reader-side day cache. Keep closed days' rows per (server, day), keyed by the build's banding inputs, and on refresh compute only [max(today − 1 day, range start), range end). Yesterday stays in the recompute set for a grace period after midnight (a few collector cadences), so late collections land. Month navigation fills misses once. In the web this is a server-side memo in the service, so every open tab shares it.
  2. Store-side closed-day rows. Materialize the per-day source aggregates once per closed day (a daily job, or a continuous aggregate per source with a daily bucket). The range read becomes rollup rows plus a raw slice for today, the shape the queries CTE already has.

Either way, a poll's cost becomes proportional to one day, not to the displayed range.

Pins:

  • result equality, cached/materialized vs today's statement, over a range that includes a late-arriving row for yesterday;
  • a live test that the second refresh of the same month reads no chunk older than yesterday (EXPLAIN chunk count, or a statement-count/rows-examined bound).

Activity

  1. added
    enhancementNew feature or request
    client-siteOwned by the client-site agents (other laptop). Local sessions never pick these up.
    on Sep 25, 2026
  2. erikdarlingdata commented on Sep 25, 2026

    @erikdarlingdata
    OwnerAuthor

    Ruling for the fix lanes (one for Darling, one for Lite):

    1. Fix shape 1, a cache on the reading side. There is no migration and no store job.
    2. A day is the day the range statement already uses. On Darling that is the UTC day, because DailySummarySql groups by date_trunc('day', collection_time) on UTC columns. The Lite lane checks which day Lite's read uses and keys its cache the same way.
    3. A day counts as closed two hours after it ends. That covers event rows that a collector picks up late, after a short outage. Today, and yesterday during those two hours, are recomputed on every refresh.
    4. The closed days of one range are cached as one block, for one hour. The first refresh computes the whole range. Later refreshes compute only the open days and join them to the cached block. After an hour, one refresh computes the whole range again. A row that arrives very late therefore shows up within the hour. Each month the calendar shows, and each rolling window, is its own block.
    5. Only reads "as of now" use the cache. A read with an explicit end time skips it.
    6. On Darling, one helper next to DailySummarySql decides what to run and joins the results. The WPF viewer keeps its own cache. The service keeps one cache for get_daily_summary_range, which the web viewer and MCP clients share.
    7. Tests: the equality test's range has a late row for yesterday inside the two hours. The cached answer must equal a fresh full statement over it. A second refresh must run the statement only over the open days. The hour limit must work.
    8. Lite first measures its month read on a 30-day fixture. If a warm read takes 100 ms or more, Lite gets the same cache. If not, the PR gives the number and Lite stays as it is.

    🤖 Generated with Claude Code

    https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3

  3. added
    in-progressActively being worked by a local session or its agents (PR open or in flight)
    on Sep 25, 2026
  4. erikdarlingdata commented on Sep 25, 2026

    @erikdarlingdata
    OwnerAuthor

    Lite half of #4232, measured per the ruling's item 8.

    Which day Lite's read groups by

    GetDailySummaryRangeAsync's SQL (Lite/Services/LocalDataService.DailySummary.cs) groups every CTE by date_trunc('day', collection_time). collection_time is a naive-UTC TIMESTAMP on every collector table. See the comment at LocalDataService.WaitStats.cs:788. ServerTab.DailySummary.cs says it directly: "The calendar buckets days in UTC." So Lite's day is the UTC day, same as Darling, with no server-offset shift. If Lite ever gets this cache, the key is the UTC day as-is.

    Fixture

    Built on SharedDuckDbFixture, the same class-scoped DuckDB fixture the neighboring Daily Summary tests (PerformanceCalendarDataTests, DailySummaryRangeToolTests) use. It is seeded with set-based INSERT ... SELECT ... FROM range(...) statements, so there are no per-row round trips, over a 30-day month:

    source cadence rows
    wait_stats 1 min, 12 wait types/collection (ScheduleManager default cadence). This is the heaviest CTE per the Darling measurement. 518,400
    cpu_utilization_stats 1 min 43,200
    query_stats 1 min, cycling 40 distinct query_hash values 43,200
    collection_log 1 min 43,200
    deadlocks sparse, spread across the month 300
    blocked_process_reports sparse 400
    dmv_blocking_snapshots sparse 400
    memory_pressure_events sparse 200
    config_alert_log sparse 150
    total 649,450

    This is heavier than any neighboring Lite DuckDB test seeds. Those seed tens to low hundreds of rows for correctness, not volume. The collection_log and query_stats rows here model only one collector's worth of cadence, which is lighter than a real multi-collector install. So this fixture is already a stress case relative to Lite's own test suite, not a favorable one.

    Measured (this machine, Debug build)

    GetDailySummaryRangeAsync(serverId, monthStart, monthStart.AddDays(30)), month fully populated, two independent runs:

    • Cold (first read after seeding, fresh query plan): 113-115 ms
    • Warm (15 repeats after 3 discarded warmups, same process/connection): min 79.4 ms, median 94.6 ms, mean 94.5 ms, max 113.7 ms

    Both runs' medians and means land under the ruling's 100 ms bar. A handful of individual samples (4 of 20 across both runs) landed just over 100 ms. The central tendency is consistently ~94-95 ms, not split around the line.

    Conclusion

    Per the ruling's item 8: warm is under 100 ms, so Lite gets no cache and no product code change. DuckDB's embedded, vectorized, single-user engine does not have the row-store page-cache-miss cost that makes Darling's PostgreSQL/TimescaleDB read expensive. That read is 4.4 s cold, and still 410 ms warm, over a comparable range. Here a full 30-day scan-and-aggregate over about 650,000 rows finishes in well under a tenth of a second.

    No PR from this lane. This closes the Lite half of #4232, with nothing further to file.

    🤖 Generated with Claude Code

    https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3

  5. erikdarlingdata commented on Sep 25, 2026

    @erikdarlingdata
    OwnerAuthor

    Closed by the watcher: delivered in PR #4307, merged to dev.

  6. removed
    in-progressActively being worked by a local session or its agents (PR open or in flight)
    on Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    client-siteOwned by the client-site agents (other laptop). Local sessions never pick these up.enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions