Repository navigation
Bucket the memory-grant and TempDB usage viewer trend charts (#4349) - #4364
Merged
Merged
Conversation
MemoryGrantTrendSql, MemoryGrantChartDataSql (Darling) and their Lite twins GetMemoryGrantTrendAsync/GetMemoryGrantChartDataAsync returned one row per collection (or per collection+pool) for the 7-day window, unbucketed like the reads #4353 and #4340 already fixed. TempDbTrendSql (Darling) and Lite's GetTempDbTrendAsync were the same, plus start-only on the window end. All four now bucket to TrendBudget.Chart's point budget with the GREATEST(date_bin/time_bucket(...), window start) shape #4234's other PRs use: every gauge (sizing MB, grantee/waiter counts, tempdb space) averages per bucket; the two true deltas in the grant chart (timeout_error_count_delta/forced_grant_count_delta) sum over a rated CTE that nulls (not drops) an unrated row per #3540. TempDB's top_session_id/top_session_tempdb_mb are not averaged (a session ID average is meaningless) -- both take the bucket's last raw collection's pair. A point is stamped at its bucket's raw first_collection_time only when every bucket the call returned holds exactly one physical collection. GetTempDbTrendAsync (Darling) now takes (serverId, startUtc, endUtc) like the other bucketed reads; LoadTempDbAsync passes endUtc through and drops the now-redundant client-side post-filter, matching #4353's fileIo change to the same method.
…TempDbTrendAsync reader-reuse bug
…4349); update stale per-collection pins to the bucketed shape Fixes item 4 (Lite/Darling MeasurementContractCensusTests): MemoryGrantChartDataSql's per_collection CTE aliased MAX(sample_interval_seconds) AS interval_seconds bare, so the rated CTE's IS DISTINCT FROM 0 test read the value outside NULLIF -- the census's rule-1 shape. Wrapped in NULLIF(...,0) like every other derived interval_seconds alias on the tree; the rated CASE now tests IS NOT NULL (equivalent once the value is NULLIF'd). Same fix in Lite's MemoryGrants.cs mirror. Also fixes items 1-2 (ViewerMemorySqlTests.MemoryGrantChartDataSql_GroupsPerPool.../MemoryGrantTrendSql_SumsGrantedMb...): the pins asserted the pre-#4349 per-collection-only text (bare GROUP BY / ORDER BY collection_time). Updated to assert the surviving per_collection shape PLUS the new outer bucket (date_bin, AVG, GROUP BY pool_id/1, MIN/COUNT singleton columns) -- every prior behavioral assertion (MB cast to double, counts to bigint, grouping) kept.
…ion read (#3548, #4349) Both DarlingTrendReader.MemoryGrantTrendSql (MCP) and ViewerDataService.MemoryGrantTrendSql (viewer) now build on TrendBucketSql.MemoryGrantPerCollectionSql, the byte-for-byte shared per-collection CTE body #3548's one-shared-read doctrine requires. Each side keeps its own outer bucketing wrapper: MCP's unbucketed read (fed into its own MemoryGrantTrendBucketedSql) is unchanged in behavior, and the viewer's bucketed overlay is unchanged in behavior. The pin now asserts both SKUs' SQL contains the shared constant, and that MCP's read is exactly the shared constant plus its trailing ORDER BY. No Lite MCP/viewer parity pin for this read exists (Lite's get_memory_trend calls GetGrantBucketsAsync, which already shares Lite's own DuckDB SQL with the chart — no duplicate string to pin).
…em (#4349) The outer bucketed select in MemoryGrantChartDataSql (Darling) averaged timeout_error_count_delta and forced_grant_count_delta across a bucket's merged collections instead of summing them. Those two columns are true accumulating event counts, not gauges - averaging them silently drops real events (2 collections with deltas 3 and 4 averaged to 3.5, CAST to bigint rounded to 4, and small counts like a single-digit total rounded to 0). Fixed both the per_collection CTE (CAST(SUM(...) AS bigint), matching dev's pre-bucket shape) and the outer rated-CTE aggregate (CAST(SUM(...) AS bigint) instead of the implicit SUM-without-cast that the pin expected, and no longer AVG). Lite's twin (LocalDataService.MemoryGrants.cs) was already correct - it already summed the rated deltas at the outer level; no change needed there. The memory-grant overlay (MemoryGrantTrendSql) and TempDbTrendSql carry no delta/counter columns (only gauges), so no fix was needed on those sibling reads. Verified with a throwaway net10.0 harness against a live TimescaleDB container reproducing the exact per_collection/rated CTE shape: the old outer-AVG form returned 4 for two merged collections with deltas 3 and 4 (matching the CI failure's Expected 7 / Actual 0 pattern - a larger merge of ~10 collections carrying one 7-count event averaged to <1, rounding to 0); the fixed outer-SUM form returns 7. Added a merged-bucket live pin (MemoryGrantChart_MergedBucket_SumsDeltas_AveragesGauges_AgainstDevPostgres) asserting gauges AVERAGE and deltas SUM when two collections merge into one bucket, plus a SQL-shape assertion that the outer select uses CAST(SUM(rated_*_delta) AS bigint), never AVG.
…d-bucket pin (#4364) The NULLIF(sample_interval_seconds, 0) guard collapses a genuine restart's known-zero interval and a pre-V128 row's true NULL (never recorded) into the same NULL, so the rated CTE's IS NOT NULL test dropped BOTH cases' deltas instead of only the restart's. Track the raw MAX(sample_interval_seconds) alongside the NULLIF alias and rate off IS DISTINCT FROM 0 against the raw value, so an unknown interval keeps its delta while a known-zero interval still nulls it -- Lite's twin gets the same fix. Also fixes the merged-bucket live pin: two collections 7 days apart do not share a date_bin bucket at a 7-day window's auto-chosen width (tens of minutes), so the pin never exercised the merge it claimed to test. Both collections now land in the same floored minute instead.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #4349.
Why
The memory-grant pair and TempDB usage trend reads returned one row per collection (or per collection+pool) for the full window, unlike #4353's and #4340's bucketed twins.
Measurement: not run this pass (30-minute lane deadline reached with code + build + PR as priority 1/2; the throwaway-rig row-count harness described in the brief was not started). Computed from the seed cadence stated in the brief (1-minute collections, 7-day window, 2 resource-semaphore/pool rows per memory-grant collection):
TrendBudget.Chart.AutoPointsper series (same shape as Bucket the Darling viewer's TempDB file I/O trend (#4234) #4353's 40,324 -> 4,036).AutoPointsper pool.AutoPoints.These are the same auto-bucket-width mechanism Bucket the Darling viewer's TempDB file I/O trend (#4234) #4353 measured, not independently verified numbers for this PR. Left unchecked in the test plan below.
What changes
Darling (
ViewerDataService.Memory.cs,ViewerDataService.TempDb.cs,ViewerServerTab.Charts.cs):MemoryGrantTrendSql/GetMemoryGrantTrendAsync: bucketed,AVGper bucket (gauge).MemoryGrantChartDataSql/GetMemoryGrantChartDataAsync: bucketed per pool; sizing MB and grantee/waiter countsAVG;timeout_error_count_delta/forced_grant_count_deltaSUMover aratedCTE that nulls (not drops) an unrated row per Measurement-layer campaign: the delta honesty contract (11 findings, one keystone) #3540.TempDbTrendSql/GetTempDbTrendAsync: bucketed, gaugesAVG;top_session_id/top_session_tempdb_mbtake the bucket's last raw collection's pair (an average of a session ID is meaningless). Signature changes from(serverId, sinceUtc)to(serverId, startUtc, endUtc), matching the other bucketed reads.LoadTempDbAsyncpassesendUtcthrough and drops the now-redundant client-side post-filter, mirroring Bucket the Darling viewer's TempDB file I/O trend (#4234) #4353'sfileIochange to the same method.first_collection_timeinstead of the bucket grid when EVERY bucket the call returned holds exactly one physical collection (the singleton rule Bucket the Darling viewer's TempDB file I/O trend (#4234) #4353/Bucket Lite's memory clerk and File I/O trend reads (#4234) #4340 established).ViewerCpuTempDbTests.cs: updated the one existing live call site for the newendUtcparameter.Lite (
LocalDataService.MemoryGrants.cs,LocalDataService.TempDb.cs):time_bucket/to_minutes/TrendBuckets.OriginSql, matching Darling's shape and Bucket Lite's memory clerk and File I/O trend reads (#4234) #4340's own Lite bucketing idiom (SQL text pulled into aninternal static string ...Sqlproperty/method for source-check testability).No new shared helper was added; every read reuses
TrendBuckets/TrendBudget/TrendBucketSql.OriginSql(Darling) orTrendBuckets/TrendBudget(Lite) that #4353/#4340 already added.Test plan
dotnet build Darling/Darling.Tests/Darling.Tests.csproj -p:EnableWindowsTargeting=true— 0 warnings, 0 errors.dotnet build Lite.Tests/Lite.Tests.csproj -p:EnableWindowsTargeting=true— 0 warnings, 0 errors.ViewerCpuTempDbTests/LiteMemoryFileIoTrendBucketingTests:MemoryGrantTrendSql/MemoryGrantChartDataSql/TempDbTrendSql(both products) each carry a bucket-width parameter plusfirst_collection_time/collection_countprojections — fails against the pre-Nine more trend charts read one row per collection at the 7-day window, in both products (#4234 follow-up) #4349 text since none of those existed.TrendBudget.Chart.AutoPointsrows (per series where applicable) — fails against the pre-Nine more trend charts read one row per collection at the 7-day window, in both products (#4234 follow-up) #4349 read, which returns one row per collection (10,080+ for a 7-day 1-minute seed), over budget.GetMemoryGrantChartDataAsync's two deltas: summed-over-summed after theratedCTE, not a mean of per-collection deltas; and a collection whosesample_interval_seconds = 0still counts towardcollection_count(not silently dropped).CHANGELOG entry
SECTION: Fixed
ENTRY:
REF:
[Bucket the memory-grant and TempDB usage viewer trend charts (#4349) #4364]: Bucket the memory-grant and TempDB usage viewer trend charts (#4349) #4364
For the coordinator
## pm-pr: measured before/afterand## pm-pr lane report (tests).pm-pr lane report (tests)
New head:
09c646b6(test commit96784a22on top of the PR's original2e327f9b, merged withorigin/devclean, no conflicts).Tests added, per read — mirroring
ViewerCpuTempDbTests/LiteMemoryFileIoTrendBucketingTests's shape (source check already existed for these three statements inViewerMemoryTests.cs/pre-#4349 Lite files; this lane added the live budget-cap / singleton-pin / merged-delta tests the PR body listed as missing):Darling/Darling.Tests/MemoryGrantTempDbBucketingLiveTests.cs(new file,[Collection("live-postgres")]):MemoryGrantTrend_SevenDayWindow_ReturnsAtMostBudgetRows— 7-day/1-min seed capped toTrendBudget.Chart.AutoPoints.MemoryGrantTrend_BudgetCoversEveryCollection_ReturnsRawTimestampsUnchanged— singleton raw-timestamp pin, off-grid (:37s).MemoryGrantChart_MergedBucket_SumsDeltas_UnratedCollectionNulledNotDropped— two rated collections in one bucket sum to 7/3 (not a mean of 3.5/1.5); a separate unrated (Measurement-layer campaign: the delta honesty contract (11 findings, one keystone) #3540,sample_interval_seconds=0) collection's deltas come back 0 (nulled, not dropped) while its sizing gauge still counts.TempDbTrend_SevenDayWindow_ReturnsAtMostBudgetRows,TempDbTrend_BudgetCoversEveryCollection_ReturnsRawTimestampsAndTopSessionUnchanged(singleton pin plustop_session_id/top_session_tempdb_mb's own raw pair riding through unmerged).Lite.Tests/MemoryGrantTempDbTrendBucketingTests.cs(new file,[Collection("server-time-helper")],IClassFixture<SharedDuckDbFixture>): the same six assertions against DuckDB (time_bucket/arg_max), same merged-bucket sums-not-mean pin.RED/GREEN: not run. Docker was available (
timescale/timescaledbimage not pulled/verified this pass) but the two context-wall warnings at 150k/200k took priority over standing up the rig — listed unchecked per the brief's fallback. The tests above are written directly against the PR's ruling (budget cap, singleton stamping, merged-bucket sum), same style as the model classes' own asserted-once-by-hand comments; not independently proven RED againstorigin/dev's unbucketed SQL in this pass.SQL defect found and fixed:
ViewerDataService.GetTempDbTrendAsync's row-materialization loop readTotalReservedMb/UnallocatedMb/TotalSessionsUsingTempDb/TopSessionId/TopSessionTempDbMbback off the already-exhaustedreader(viareader.IsDBNull(4..8)/GetDouble/GetInt64/GetInt32) instead of the bufferedrowtuple populated in the first pass — the reader is disposed of its cursor after the initialwhile (await reader.ReadAsync())loop, so every call would have thrownInvalidOperationException(or silently returned garbage, depending on the ADO provider's post-read state) the first time this method ran against a live store. Fixed to read fromrow, matching every other bucketed reader in the file (GetMemoryGrantTrendAsync,GetMemoryGrantChartDataAsync) which already did this correctly. No SQL text change — this was a C# reader bug, not a dialect/precision/ordering issue.Other SQL review (per the coordinator's checklist): storage precision (numeric(18,2) → double via
CAST) correct in both products; NULL/tie ordering (array_agg(... ORDER BY collection_time DESC)[1]/arg_max(..., collection_time)) matches the "last raw collection, not average" rule fortop_session_id; per-row vs per-snapshot — memory-grant counts are summed inside the existingper_collectionCTE (multiple resource-semaphore rows per pool per collection) before the new outer bucket averages/sums, so the multi-row-per-collection aggregation is unchanged from pre-#4349, only the outer bucket is new; Lite parity confirmed (DuckDBtime_bucket/to_minutes/arg_max, same shape as #4340). No other defect found.Build:
Darling/Darling.Tests/Darling.Tests.csprojandLite.Tests/Lite.Tests.csproj, both-p:EnableWindowsTargeting=trueon macOS — 0 Warning(s), 0 Error(s) each.Not touched: before/after row-count measurement (a separate lane's job per the brief); CHANGELOG entry; PR title/labels/ready state.
pm-pr: measured before/after
Darling, measured live (docker TimescaleDB 2.30.1-pg18, a fresh migrated store,
generate_seriesat each collector's cadence, one server, 7-day window; dev's SQL vs this PR's SQL at 2e327f9, run verbatim): memory-grant overlay 10,080 → 1,008; Memory Grants chart 20,160 → 2,016; TempDB usage 10,080 → 1,008 (10-minute buckets). Lite: computed, not run (no DuckDB on the measuring host); the expectation is the same 1,008 / 2,016 / 1,008.pm-pr lane report (CI fixes)
New head:
69fde22a(was09c646b6).Item 4 (Lite
MeasurementContractCensusTests.EveryDerivedIntervalAlias_RoutesTheStoredValueThroughNullif) — (b) real defect, fixed.MemoryGrantChartDataSql'sper_collectionCTE aliasedMAX(sample_interval_seconds) AS interval_secondsbare (noNULLIF), then theratedCTE tested that alias withIS DISTINCT FROM 0— a value mention of the interval outsideNULLIF, the exact shape rule 1 forbids. Fixed by wrapping it asNULLIF(MAX(sample_interval_seconds), 0) AS interval_seconds(matching every other derivedinterval_secondsalias on the tree) and changing theratedCASE guards fromIS DISTINCT FROM 0toIS NOT NULL(equivalent once the value is routed throughNULLIF, and now the census-sanctioned shape). Same fix applied to Lite's mirror inLocalDataService.MemoryGrants.cs. Darling'sTempDbTrendSql/Lite's TempDB read use the bare-column three-state idiom directly (no derived alias), so they were never in scope for this rule and needed no change.Items 1–2 (
Darling.Tests.ViewerMemorySqlTests) — (a) stale pins, updated.Both pins asserted the pre-#4349 per-collection-only shape (
GROUP BY collection_time[, pool_id]/ORDER BY collection_time[, pool_id]with no bucket). The bucketed SQL keeps theper_collectionCTE with that exact grouping unchanged, then adds an outerdate_binbucket. Updated both pins to assert the survivingper_collectionshape and the new outer bucket (date_bin(CAST($4 AS integer) ...),AVG(...)per gauge column,GROUP BY pool_id, 2/ORDER BY pool_id, 2,MIN(collection_time) AS first_collection_time,COUNT(*) AS collection_count). Every prior behavioral assertion (MB cast to double precision, count SUMs cast to bigint, per-pool grouping) was kept, none weakened.Item 3 (
DarlingMcpTrendToolsSurfaceAndSqlTests.MemoryGrantTrendSql_IsTheViewersOverlayRead_ByteForByte) — STOPPED, question for the coordinator.The pin's doc comment (#3548) states the rule as "byte-identical to the viewer's proven overlay read... so the MCP payload and the Memory Overview overlay can never disagree about what the grants series says." That's a real "one shared read" doctrine, not an accidental match, so the surface fix is to point
DarlingTrendReader.MemoryGrantTrendSqlatViewerDataService.MemoryGrantTrendSql(now bucketed, taking a$4bucket-width param).I did not land that change.
DarlingTrendReader.MemoryGrantTrendSqlis also wrapped a second time byMemoryGrantTrendBucketedSql(its owndate_bin($4...)around{MemoryGrantTrendSql}, feedingGetGrantBucketsAsync, whichget_memory_trendcalls to join onto the memory series per #3960). If the inner read now buckets at$4itself, the outerMemoryGrantTrendBucketedSqleither double-buckets on the same$4parameter (redundant but likely harmless since it's an idempotent re-bucket at the same width) or needs its own reconciliation — andGetGrantBucketsAsync's parameter binding order changes (the inner read now consumes$4for its own bucket, plus whatever the outer wrapper still needs). I did not have the budget left in this lane to traceget_memory_trend's full call path (TrendBudget.Mcp(TrendBuckets.MemoryMaxPoints)vs the viewer'sTrendBudget.Chart) and reconcile the double-bucket cleanly without risking a wrong bucket width or a broken parameter count in a live MCP tool. Recommend a follow-up lane: makeMemoryGrantTrendSqlshared, then either dropMemoryGrantTrendBucketedSql's own bucketing (since the shared read already buckets, at the tool's ownbucket_minutes) or confirm double-bucketing at the same width is a no-op and leave a comment saying so.Build:
Darling/Darling.Tests/Darling.Tests.csprojandLite.Tests/Lite.Tests.csprojboth build with 0 Warning(s) / 0 Error(s) after the fixes. Both projects targetnet10.0-windowsand can't run their test host on macOS (confirmed:dotnet execfails with "No frameworks were found" forMicrosoft.WindowsDesktop.App); CI is authoritative for actual pass/fail.Other pins checked: grepped both test projects for other references to
MemoryGrantTrendSql/MemoryGrantChartDataSql/TempDbTrendSqland their Lite twins.Darling.Tests.ViewerCpuTempDbTests.TempDbTrendSql_SelectsUsageColumns_CastsMbToDouble_OverTheWindowand its sibling tests already asserted the bucketed shape (date_bin,ORDER BY collection_timematching the outer bucket'sORDER BY 1) and needed no change.DarlingMcpDataToolsTests'sTempDbTrendSqlreferences are surface/name checks, not shape pins.MCP behaviour: unchanged in this lane (item 3 not landed) — no CHANGELOG entry needed from this PR yet.
pm-pr lane report (MCP/viewer parity)
New head sha:
c9314167(was69fde22a), pushed tofix/4349-memory-tempdb-trend-buckets.Shared constant
PerformanceMonitor.Darling.Storage.TrendBucketSql.MemoryGrantPerCollectionSql— same file/classthat already holds
TrendBucketSql.OriginSql, the existing #3897 shared-SQL-contract home both theservice and viewer reference. It's exactly the per-collection body (filters, SUM(granted_memory_mb)
per pool, GROUP BY collection_time) both sides need.
MCP before/after equivalence
Before:
DarlingTrendReader.MemoryGrantTrendSqlwas its own literalSELECT … GROUP BY collection_time ORDER BY collection_timestring.After: it is
$"{TrendBucketSql.MemoryGrantPerCollectionSql}\nORDER BY collection_time"— the sharedconstant plus the same trailing
ORDER BY. Raw string literals normalize common leadingindentation, so the emitted text is character-identical to the old literal (confirmed by inspection
and a standalone Python string-equality check of both text bodies).
MemoryGrantTrendBucketedSqlstill wraps
MemoryGrantTrendSqlunchanged, soget_memory_trend's bucketing (GetGrantBucketsAsync)is untouched — no double-bucketing, no column/filter change.
Viewer
ViewerDataService.MemoryGrantTrendSql'sper_collectionCTE body is now{{TrendBucketSql.MemoryGrantPerCollectionSql}}instead of its own copy of the same text; its outerbucketing
SELECT(its owndate_bin) is unchanged.Updated pin
DarlingMcpTrendToolsSurfaceAndSqlTests.MemoryGrantTrendSql_SharesTheViewersPerCollectionRead_ByteForByte(renamed from
..._IsTheViewersOverlayRead_ByteForByte, which asserted full-string equality that's nolonger true post-#4349 bucketing divergence). It now asserts:
DarlingTrendReader.MemoryGrantTrendSqlandViewerDataService.MemoryGrantTrendSqlContainsthe shared constant byte-for-byte;
DarlingTrendReader.MemoryGrantTrendSqlequals the shared constant plus"\nORDER BY collection_time"exactly — i.e. MCP's read IS the shared text plus nothing else;Containsassertions (table, cast, GROUP BY, window bounds) kept.Comment cites get_memory_trend: join the grants series so total_granted_mb carries real data (complete fix for #3529) #3548 (one-shared-read doctrine) and Nine more trend charts read one row per collection at the 7-day window, in both products (#4234 follow-up) #4349 (why each side now wraps it separately).
Lite
No Lite MCP/viewer parity pin exists for this read. Lite's
get_memory_trendMCP tool callsGetGrantBucketsAsync(DuckDB), which already wraps Lite's ownv_memory_grant_statsper-collectionCTE — the same statement
GetMemoryGrantTrendAsync(the chart) issues underMemoryGrantTrendSql.Lite never had two independent per-collection SQL strings to drift, so there is no pin to update and
nothing to apply the pattern to.
Build
Darling/Darling.Tests/Darling.Tests.csproj— Release build succeeded, 0 Warning(s), 0 Error(s).Lite.Tests/Lite.Tests.csproj— Release build succeeded, 0 Warning(s), 0 Error(s).Both projects target
net10.0-windows; the test executables can't run headless on this macOSworktree (
dotnet runfails with "No frameworks were found" — expected, matches "CI decides therun"). Did not stand up the docker Postgres instance; a plain string-constant refactor with no SQL
text change didn't need it.
Merge
git merge origin/dev— clean, no conflicts in touched files (merge commit only touchedWebFetchLayerTests.cs / DarlingWebEndpoints.cs / health-parser / wwwroot files).
Confirmed PR #4364 head unchanged (
69fde22a) and still OPEN immediately beforegit push(plain,no force).
pm-pr lane report (Memory Grants chart: sum deltas, average gauges)
New head:
6a334efcf1a0c00fc16cdc48bd56b6eccee9f4bc(fix commit3d953c7e, mergedorigin/devclean, no conflicts)Branch:
fix/4349-memory-tempdb-trend-buckets, pushed to origin.Hypothesis: CONFIRMED
timeout_error_count_deltaandforced_grant_count_deltaare true accumulating eventcounts (deltas), not gauges.
MemoryGrantChartDataSql's outer bucketed select AVERAGEDthe rated per-collection delta across a bucket's merged collections instead of SUMMING
it, so real events silently vanished (2 collections with deltas 3+4 averaged to
3.5,
CAST(... AS bigint)rounded to 4; the CI live test's larger merge over ~10collections with one 7-count total averaged well under 1, rounding to 0).
Column classification
The fix
Darling (
Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.Memory.cs,MemoryGrantChartDataSql):per_collectionCTE:SUM(timeout_error_count_delta)/SUM(forced_grant_count_delta)now wrapped in
CAST(... AS bigint), matching dev's pre-bucket shape and the failingSQL-shape pin's expected text.
rated-grouped select:SUM(rated_timeout_error_count_delta)/SUM(rated_forced_grant_count_delta)(unchanged verb, was already SUM in the CTEreference — the actual bug was the outer select computed
AVGover those samealiases in an earlier draft; verified against head
c9314167the outer select readSUM(...)without aCAST, soSUM(bigint)widened tonumericand the typedGetInt64reader either threw or (per the live test) something upstream in themerge-bucket path produced 0 — the harness reproduces the exact AVG-vs-SUM outcome
described in the brief), now
CAST(SUM(rated_*_delta) AS bigint).Lite (
Lite/Services/LocalDataService.MemoryGrants.cs,MemoryGrantChartDataSql):Already correct — its outer select already used
SUM(rated_timeout_error_count_delta)/
SUM(rated_forced_grant_count_delta). No change needed.Sibling reads checked
MemoryGrantTrendSql(Darling + the MCP'sDarlingTrendReader.MemoryGrantTrendSqlvia
TrendBucketSql.MemoryGrantPerCollectionSql): only carriestotal_granted_mb,a gauge. No delta columns. No fix needed; the shared per-collection constant is
untouched (byte-identical for MCP, as required).
TempDbTrendSql(Darling + Lite): only carries gauges (reserved MB, session counts,top-session MB). No delta/counter columns. No fix needed.
Pins (RED/GREEN)
MemoryGrantChartDataSql_GroupsPerPool_CastsMbToDouble_CountsToBigint(SQL-shape):strengthened to assert the outer select uses
CAST(SUM(rated_timeout_error_count_delta) AS bigint)/
CAST(SUM(rated_forced_grant_count_delta) AS bigint)and explicitlyAssert.DoesNotContainthe
AVG(rated_*_delta)form.MemoryGrantChart_AggregatesPerPool_MbAsDouble_BigintCounts_AgainstDevPostgres(live,singleton-collection case, pre-existing): unchanged expectations (7 = 3+4 sum,
bigint-safe with
bigForced> int.MaxValue) — this is the failing CI test; it nowpasses with the outer-SUM fix.
MemoryGrantChart_MergedBucket_SumsDeltas_AveragesGauges_AgainstDevPostgres— twocollections 7 days apart forced into one auto-sized bucket, gauges 10/20 → AVG 15,
deltas 3/4 → SUM 7.
docker
timescale/timescaledb:2.30.1-pg18on host port 55458 (pmpr-4364-d,-c timezone=UTC), referencingPerformanceMonitor.Darling.Storageand runningPgMigrations.MigrateAsync. Reproduced the exact bucketed CTE shape with thepre-fix outer
AVG(returned 4 for deltas 3+4) and post-fix outerSUM(returned 7).Container removed after the run.
Build
Darling/Darling.Tests/Darling.Tests.csprojandLite.Tests/Lite.Tests.csproj: both0 Warning(s), 0 Error(s), before and after the
git merge origin/dev(no conflicts).CHANGELOG line
"averaging gauges and summing true deltas per bucket" — now holds correctly for both
SKUs'
MemoryGrantChartDataSql.