Skip to content

Cache the I/O-latency and CPU baselines by UTC day, not the hour (#4248) - #4291

Merged
erikdarlingdata merged 2 commits into
devfrom
fix/4248-daily-raw-baselines
Sep 25, 2026
Merged

erikdarlingdata merged 2 commits into
devfrom
fix/4248-daily-raw-baselines

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Closes #4248.

Why

PgBaselineProvider's IoLatency arm reads the raw file_io_stats hypertable at per-file grain over the full 30-day baseline window. The cache key was the analysis hour. So every server recomputed that 30-day median/MAD scan every hour, even though a 30-day baseline moves by about 1/720 an hour. The issue measured about 50 MB of temp spill per call on that arm alone. That was 13.5 of 22.1 GB of one store's total temp writes since its stats reset.

Confirmed on the rig: Cpu reads its own raw table (cpu_utilization_stats) the same way. The code already documents this split, in the #1743 comment on GetBaselineQuery.

Of the nine RobustTierScaffold arms, only Cpu and IoLatency read a CREATE TABLE hypertable directly. The other seven arms read a _baseline materialized view that holds one row per collection, which is far fewer rows for the same 30 days. They are BatchRequests, WaitStats, SessionCount, QueryDuration, Memory, WaitMsPerSec and BlockingPerMinute. The two event arms (Blocking, Deadlock) bypass the scaffold entirely. So Cpu and IoLatency are the full set of arms this ruling applies to in PgBaselineProvider.cs.

What changes

PgTargetBaselineProvider.cs (PostgreSQL targets) is left alone. It uses the same cache, and every one of its arms reads a raw hypertable. Its metric names are not in IsDailyCacheMetric's set, so it does not change here. It has keyed arms (one series per statement), and KeyedBaselineCacheWarnCount's cardinality warning assumes hourly turnover. So a day-long lifetime there needs its own ruling. #4298 tracks it.

Pins added

  • Darling, live (DarlingAnomalyBaselineTests.cs): EndToEnd_IoLatencyArm_TwoCallsHoursApartOnOneDay_ShareOneCompute_NextDayRecomputes_AgainstDevPostgres. It seeds file_io_stats, then calls GetBaselineAsync at hour 1 and hour 20 of the same UTC day. It counts 0 new reads, with CommandCapture, the existing Every scheduled analysis pass recomputes every 30-day baseline: a fresh DarlingAnalysisService per pass never reuses (or shares with MCP) the bucket cache #3941 technique. A call the next UTC day counts 1 new read.
  • Darling, updated: PgTargetClockTests.cs's literal source-text assertion now expects RoundedKeyTime(metricName, analysisTime), replacing the old RoundedHour(analysisTime).
  • Lite: DailyCacheEntry_AnswersForItsWholeUtcDay_NotJustTheTtl_AndStopsAtMidnight pins the day-long liveness and the midnight cutoff directly against BaselineCache, with a synthetic entry and no DB seed needed. ASharedEntry_IsTheFreshAnswer_AndReadsNothing exercises MetricNames.Cpu directly, and its "next analysis hour is another window" case needed a correction: that is no longer true for Cpu under this change. The comment and assertion now say so, instead of passing on a check that had gone silently weak.

Lite's full suite ran clean, and Darling's four failures are in code this PR does not touch (below). So I did not add an automated pg_stat_statements.temp_blks_written live test. Gating it correctly on whether CI's darling-pg job preloads pg_stat_statements needed more checking than this lane had time for. I measured it by hand on the rig instead:

before (hourly key) after (daily key)
IoLatency, 3 calls same UTC day, over an hour apart 3 computes, 3 potential spills 1 compute
IoLatency, 1 call the next UTC day recomputes (expected) recomputes (expected)

Test plan

  • Darling.Tests: 0 Warning(s), 0 Error(s).
  • Lite.Tests: 0 Warning(s), 0 Error(s).
  • Full Darling.Tests.exe run once, after git merge origin/dev and a dropped and recreated darlingtest: 14004 total, 0 errors, 4 failed, 48 skipped, 1 not run.
    • All 4 failures are EXPLAIN-plan-shape and chunk-pruning tests on collection_log and wait-rate tiles (ForcePlanFailuresAccessPathTests, WaitRateTileReadCountLiveTests, ServerListAndSummaryPlanShapeTests, CaptureDownChunkOrderTests). None relate to PgBaselineProvider or BaselineCache, the only production files this PR touches.
    • I re-ran each alone on the same database. 2 of 4 passed, which points to interference between tests on the shared darlingtest database. The other 2 (ServerListAndSummaryPlanShapeTests, CaptureDownChunkOrderTests) still failed alone on that rig. This diff touches none of the code they test.
  • Full Lite.Tests.exe run once: 5373 total, 0 errors, 0 failed, 0 skipped.
  • Diagnosis confirmed on the rig with the measured table above, not just asserted from the issue's production numbers.
  • GitHub Actions CI on 8a1f0ee9: build, Darling PostgreSQL tests and Lite tests all pass.

CHANGELOG entry

SECTION: Fixed
ENTRY:

🤖 Generated with Claude Code

https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ

erikdarlingdata and others added 2 commits September 25, 2026 10:23
…TC day (#4248)

IoLatency (file_io_stats, per-file rows) and Cpu (cpu_utilization_stats,
per-sample rows) are the two PgBaselineProvider arms that read a RAW
hypertable at full grain over the 30-day baseline window instead of a
pre-aggregated "_baseline" materialized view. Keying them hourly meant
recomputing a 30-day median/MAD over that full row set every hour, even
though the baseline itself moves by about 1/720 an hour; IoLatency alone
measured ~50 MB of temp spill per call in production.

Both arms now key on the UTC day: RoundedKeyTime(metricName, analysisTime)
picks RoundedDay for these two and keeps RoundedHour for every other arm,
and it is the one seam both the cache key and the compute's window end
read, so #3941's "one key, one set of rows" invariant holds at either
grain. A successful daily compute gets a CachedBaseline.FreshUntilUtc
(RealTime + 24h) so a lookup never falls back to the 1-hour CacheTtl for
these two; a FAILED compute keeps the hourly retry regardless of metric.
BaselineCache (the process-wide shared tier) and its sweep read the same
PgBaselineProvider.IsFresh, so they honor the same lifetime.

PgTargetBaselineProvider (PostgreSQL-monitored targets) inherits this
mechanism unchanged: its metric names are not in IsDailyCacheMetric's set,
even though its own doc comment notes every one of its arms also reads a
raw hypertable. Left alone deliberately — its keyed arms raise a lifetime
question against KeyedBaselineCacheWarnCount's hourly-turnover calibration
that nobody has ruled on. Flagged in the PR for the coordinator.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
Same fix as Darling's PgBaselineProvider, twinned: RoundedKeyTime picks
RoundedDay for the two arms that read a RAW table (v_cpu_utilization_stats,
v_file_io_stats) at full grain over the 30-day window, RoundedHour for
every other arm, and both the cache key and the compute's window end read
that one seam. A successful daily compute gets CachedBaseline.FreshUntilUtc
(RealTime + 24h); a failed one keeps the hourly retry. BaselineCache (the
per-store shared tier) and its sweep read the same BaselineProvider.IsFresh.

Updated SharedBaselineCacheTests' ASharedEntry_IsTheFreshAnswer_AndReadsNothing,
which exercises MetricNames.Cpu directly: its "next analysis hour is another
window" case is no longer true for this metric (same UTC day, so still one
shared entry) — the comment and assertion now say so, and a new
DailyCacheEntry_AnswersForItsWholeUtcDay_NotJustTheTtl_AndStopsAtMidnight pins
the day-long liveness and the midnight cutoff directly against BaselineCache,
independent of any seeded data.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review September 25, 2026 15:12
@erikdarlingdata
erikdarlingdata merged commit ff376dc into dev Sep 25, 2026
16 of 17 checks passed
@erikdarlingdata
erikdarlingdata deleted the fix/4248-daily-raw-baselines branch September 25, 2026 15:12
erikdarlingdata added a commit that referenced this pull request Sep 25, 2026
Four tests next to #4291's/#4248's: a PostgreSQL-target arm's key is the
UTC day and two analysis times in the same day round to one entry; a
keyed arm (pg_statement_mean_ms) walked across a UTC day boundary holds
at most two shared-tier entries per member; a failed compute still
retries within the hour rather than inheriting the day-long lifetime a
success gets; and the SQL Server arms' keys (Cpu/IoLatency daily, the
rest hourly) are unchanged. The first two fail against the pre-fix
override body (proven by hand, reverting it, before committing).

Also bumps PgTargetClockTests' seam census from three to four now that
IsDailyCacheArm joins ResolveBaselineQuery, ReadServerClockAsync and
ResolveKeyedBaselineQuery as a redeclared seam.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
erikdarlingdata added a commit that referenced this pull request Sep 25, 2026
…) (#4324)

* Make the daily-cache-arm decision an overridable seam (#4298)

PgBaselineProvider.IsDailyCacheMetric only ever named Cpu and IoLatency,
so PgTargetBaselineProvider's 15 arms (13 unkeyed, 2 keyed) all keyed and
recomputed hourly even though every one reads a raw PostgreSQL-target
hypertable at full grain over the 30-day window - lane B2a measured up to
2.45s and 262MB of temp per compute (pg_statement_mean_ms keyed, worst).

Add IsDailyCacheArm as a fourth protected virtual seam beside
ResolveBaselineQuery, ReadServerClockAsync and ResolveKeyedBaselineQuery,
defaulting to the existing IsDailyCacheMetric rule. RoundedKeyTime (the
one seam both the cache key and the compute's window end read) now reads
it instead, so it moves from static to instance. PgTargetBaselineProvider
overrides it to return true unconditionally for all its arms. SQL Server
arms are unaffected: IsDailyCacheMetric's own body, and the base class's
answer, are untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3

* Pin the day-keyed PgTargetBaselineProvider cache (#4298)

Four tests next to #4291's/#4248's: a PostgreSQL-target arm's key is the
UTC day and two analysis times in the same day round to one entry; a
keyed arm (pg_statement_mean_ms) walked across a UTC day boundary holds
at most two shared-tier entries per member; a failed compute still
retries within the hour rather than inheriting the day-long lifetime a
success gets; and the SQL Server arms' keys (Cpu/IoLatency daily, the
rest hourly) are unchanged. The first two fail against the pre-fix
override body (proven by hand, reverting it, before committing).

Also bumps PgTargetClockTests' seam census from three to four now that
IsDailyCacheArm joins ResolveBaselineQuery, ReadServerClockAsync and
ResolveKeyedBaselineQuery as a redeclared seam.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3

* Fix five daily-cache-tier test premises after #4298 (CI for #4324)

All five were verdict (a): the daily cache tier changed a window end or a
snapshot-window relationship the test's fixture assumed, not a product bug.

- PgTargetBetweenWavesV3Tests: the doc pin still expected "three seams";
  #4298 rewrote the doc to "four" when it added IsDailyCacheArm.
- SharedBaselineCacheLiveTests (two tests): PgTps and PgStatementShare are
  daily-cache arms now, so the next HOUR is a cache hit (not a recompute),
  and an entry stays live until FreshUntilUtc (a rolling 24h), not CacheTtl.
- PgTargetClockLiveTests: the daily window now ends at the UTC day boundary
  at or before windowStart, not at the analysis hour, so the session-scoped
  snapshot needed replanting at or before that boundary to still land
  "inside the window" the snapshot rule is proving.
- PgTargetSampledWaitLiveTests: the young server's six hours of history sit
  entirely after the new day boundary, so the standalone baseline lookup and
  the detector's own internal one (same provider, same analysisTime) both
  saw zero samples. Added two isolated quiet hours anchored at the boundary,
  on one calendar date, so the bucket has samples without crossing the
  three-distinct-day trustworthy floor.

Darling.Tests: 0 warnings, 0 errors. 135 tests run across the five failing
classes plus DarlingAnomalyBaselineTests, PgBaselineProviderKeyed(Live)Tests
and DocCommentHygieneTests — 0 failed, 0 skipped.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant