Skip to content

Bucket the memory-grant and TempDB usage viewer trend charts (#4349) - #4364

Merged
erikdarlingdata merged 11 commits into
devfrom
fix/4349-memory-tempdb-trend-buckets
Sep 26, 2026
Merged

erikdarlingdata merged 11 commits into
devfrom
fix/4349-memory-tempdb-trend-buckets

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Refs #4349.

Why

The memory-grant pair and TempDB usage trend reads returned one row per collection (or per collection+pool) for the full window, unlike #4353's and #4340's bucketed twins.

Measurement: not run this pass (30-minute lane deadline reached with code + build + PR as priority 1/2; the throwaway-rig row-count harness described in the brief was not started). Computed from the seed cadence stated in the brief (1-minute collections, 7-day window, 2 resource-semaphore/pool rows per memory-grant collection):

What changes

Darling (ViewerDataService.Memory.cs, ViewerDataService.TempDb.cs, ViewerServerTab.Charts.cs):

Lite (LocalDataService.MemoryGrants.cs, LocalDataService.TempDb.cs):

No new shared helper was added; every read reuses TrendBuckets/TrendBudget/TrendBucketSql.OriginSql (Darling) or TrendBuckets/TrendBudget (Lite) that #4353/#4340 already added.

Test plan

  • dotnet build Darling/Darling.Tests/Darling.Tests.csproj -p:EnableWindowsTargeting=true — 0 warnings, 0 errors.
  • dotnet build Lite.Tests/Lite.Tests.csproj -p:EnableWindowsTargeting=true — 0 warnings, 0 errors.
  • Before/after row-count measurement (not run this pass; see Why).
  • New bucketing regression tests (not added this pass — ran out of the 30-minute budget after the fix + build + PR priority). Would need, mirroring ViewerCpuTempDbTests/LiteMemoryFileIoTrendBucketingTests:
    • A source-check test pinning MemoryGrantTrendSql/MemoryGrantChartDataSql/TempDbTrendSql (both products) each carry a bucket-width parameter plus first_collection_time/collection_count projections — fails against the pre-Nine more trend charts read one row per collection at the 7-day window, in both products (#4234 follow-up) #4349 text since none of those existed.
    • A live DuckDB/Postgres 7-day one-per-minute seed test asserting each read returns at most TrendBudget.Chart.AutoPoints rows (per series where applicable) — fails against the pre-Nine more trend charts read one row per collection at the 7-day window, in both products (#4234 follow-up) #4349 read, which returns one row per collection (10,080+ for a 7-day 1-minute seed), over budget.
    • A point-equality test: a handful of collections inside one auto-picked bucket width come back at their own raw timestamps unchanged, not floored to the bucket grid (the singleton-stamping rule).
    • A merged-bucket test for GetMemoryGrantChartDataAsync's two deltas: summed-over-summed after the rated CTE, not a mean of per-collection deltas; and a collection whose sample_interval_seconds = 0 still counts toward collection_count (not silently dropped).
  • Full Windows Darling.Tests / Lite.Tests suite (CI decides; not runnable on macOS).

CHANGELOG entry

SECTION: Fixed
ENTRY:

For the coordinator

pm-pr lane report (tests)

New head: 09c646b6 (test commit 96784a22 on top of the PR's original 2e327f9b, merged with origin/dev clean, no conflicts).

Tests added, per read — mirroring ViewerCpuTempDbTests/LiteMemoryFileIoTrendBucketingTests's shape (source check already existed for these three statements in ViewerMemoryTests.cs/pre-#4349 Lite files; this lane added the live budget-cap / singleton-pin / merged-delta tests the PR body listed as missing):

  • Darling Darling/Darling.Tests/MemoryGrantTempDbBucketingLiveTests.cs (new file, [Collection("live-postgres")]):
    • MemoryGrantTrend_SevenDayWindow_ReturnsAtMostBudgetRows — 7-day/1-min seed capped to TrendBudget.Chart.AutoPoints.
    • MemoryGrantTrend_BudgetCoversEveryCollection_ReturnsRawTimestampsUnchanged — singleton raw-timestamp pin, off-grid (:37s).
    • MemoryGrantChart_MergedBucket_SumsDeltas_UnratedCollectionNulledNotDropped — two rated collections in one bucket sum to 7/3 (not a mean of 3.5/1.5); a separate unrated (Measurement-layer campaign: the delta honesty contract (11 findings, one keystone) #3540, sample_interval_seconds=0) collection's deltas come back 0 (nulled, not dropped) while its sizing gauge still counts.
    • TempDbTrend_SevenDayWindow_ReturnsAtMostBudgetRows, TempDbTrend_BudgetCoversEveryCollection_ReturnsRawTimestampsAndTopSessionUnchanged (singleton pin plus top_session_id/top_session_tempdb_mb's own raw pair riding through unmerged).
  • Lite Lite.Tests/MemoryGrantTempDbTrendBucketingTests.cs (new file, [Collection("server-time-helper")], IClassFixture<SharedDuckDbFixture>): the same six assertions against DuckDB (time_bucket/arg_max), same merged-bucket sums-not-mean pin.

RED/GREEN: not run. Docker was available (timescale/timescaledb image not pulled/verified this pass) but the two context-wall warnings at 150k/200k took priority over standing up the rig — listed unchecked per the brief's fallback. The tests above are written directly against the PR's ruling (budget cap, singleton stamping, merged-bucket sum), same style as the model classes' own asserted-once-by-hand comments; not independently proven RED against origin/dev's unbucketed SQL in this pass.

SQL defect found and fixed: ViewerDataService.GetTempDbTrendAsync's row-materialization loop read TotalReservedMb/UnallocatedMb/TotalSessionsUsingTempDb/TopSessionId/TopSessionTempDbMb back off the already-exhausted reader (via reader.IsDBNull(4..8)/GetDouble/GetInt64/GetInt32) instead of the buffered row tuple populated in the first pass — the reader is disposed of its cursor after the initial while (await reader.ReadAsync()) loop, so every call would have thrown InvalidOperationException (or silently returned garbage, depending on the ADO provider's post-read state) the first time this method ran against a live store. Fixed to read from row, matching every other bucketed reader in the file (GetMemoryGrantTrendAsync, GetMemoryGrantChartDataAsync) which already did this correctly. No SQL text change — this was a C# reader bug, not a dialect/precision/ordering issue.

Other SQL review (per the coordinator's checklist): storage precision (numeric(18,2) → double via CAST) correct in both products; NULL/tie ordering (array_agg(... ORDER BY collection_time DESC)[1] / arg_max(..., collection_time)) matches the "last raw collection, not average" rule for top_session_id; per-row vs per-snapshot — memory-grant counts are summed inside the existing per_collection CTE (multiple resource-semaphore rows per pool per collection) before the new outer bucket averages/sums, so the multi-row-per-collection aggregation is unchanged from pre-#4349, only the outer bucket is new; Lite parity confirmed (DuckDB time_bucket/to_minutes/arg_max, same shape as #4340). No other defect found.

Build: Darling/Darling.Tests/Darling.Tests.csproj and Lite.Tests/Lite.Tests.csproj, both -p:EnableWindowsTargeting=true on macOS — 0 Warning(s), 0 Error(s) each.

Not touched: before/after row-count measurement (a separate lane's job per the brief); CHANGELOG entry; PR title/labels/ready state.

pm-pr: measured before/after

Darling, measured live (docker TimescaleDB 2.30.1-pg18, a fresh migrated store, generate_series at each collector's cadence, one server, 7-day window; dev's SQL vs this PR's SQL at 2e327f9, run verbatim): memory-grant overlay 10,080 → 1,008; Memory Grants chart 20,160 → 2,016; TempDB usage 10,080 → 1,008 (10-minute buckets). Lite: computed, not run (no DuckDB on the measuring host); the expectation is the same 1,008 / 2,016 / 1,008.

pm-pr lane report (CI fixes)

New head: 69fde22a (was 09c646b6).

Item 4 (Lite MeasurementContractCensusTests.EveryDerivedIntervalAlias_RoutesTheStoredValueThroughNullif) — (b) real defect, fixed.
MemoryGrantChartDataSql's per_collection CTE aliased MAX(sample_interval_seconds) AS interval_seconds bare (no NULLIF), then the rated CTE tested that alias with IS DISTINCT FROM 0 — a value mention of the interval outside NULLIF, the exact shape rule 1 forbids. Fixed by wrapping it as NULLIF(MAX(sample_interval_seconds), 0) AS interval_seconds (matching every other derived interval_seconds alias on the tree) and changing the rated CASE guards from IS DISTINCT FROM 0 to IS NOT NULL (equivalent once the value is routed through NULLIF, and now the census-sanctioned shape). Same fix applied to Lite's mirror in LocalDataService.MemoryGrants.cs. Darling's TempDbTrendSql/Lite's TempDB read use the bare-column three-state idiom directly (no derived alias), so they were never in scope for this rule and needed no change.

Items 1–2 (Darling.Tests.ViewerMemorySqlTests) — (a) stale pins, updated.
Both pins asserted the pre-#4349 per-collection-only shape (GROUP BY collection_time[, pool_id] / ORDER BY collection_time[, pool_id] with no bucket). The bucketed SQL keeps the per_collection CTE with that exact grouping unchanged, then adds an outer date_bin bucket. Updated both pins to assert the surviving per_collection shape and the new outer bucket (date_bin(CAST($4 AS integer) ...), AVG(...) per gauge column, GROUP BY pool_id, 2 / ORDER BY pool_id, 2, MIN(collection_time) AS first_collection_time, COUNT(*) AS collection_count). Every prior behavioral assertion (MB cast to double precision, count SUMs cast to bigint, per-pool grouping) was kept, none weakened.

Item 3 (DarlingMcpTrendToolsSurfaceAndSqlTests.MemoryGrantTrendSql_IsTheViewersOverlayRead_ByteForByte) — STOPPED, question for the coordinator.
The pin's doc comment (#3548) states the rule as "byte-identical to the viewer's proven overlay read... so the MCP payload and the Memory Overview overlay can never disagree about what the grants series says." That's a real "one shared read" doctrine, not an accidental match, so the surface fix is to point DarlingTrendReader.MemoryGrantTrendSql at ViewerDataService.MemoryGrantTrendSql (now bucketed, taking a $4 bucket-width param).

I did not land that change. DarlingTrendReader.MemoryGrantTrendSql is also wrapped a second time by MemoryGrantTrendBucketedSql (its own date_bin($4...) around {MemoryGrantTrendSql}, feeding GetGrantBucketsAsync, which get_memory_trend calls to join onto the memory series per #3960). If the inner read now buckets at $4 itself, the outer MemoryGrantTrendBucketedSql either double-buckets on the same $4 parameter (redundant but likely harmless since it's an idempotent re-bucket at the same width) or needs its own reconciliation — and GetGrantBucketsAsync's parameter binding order changes (the inner read now consumes $4 for its own bucket, plus whatever the outer wrapper still needs). I did not have the budget left in this lane to trace get_memory_trend's full call path (TrendBudget.Mcp(TrendBuckets.MemoryMaxPoints) vs the viewer's TrendBudget.Chart) and reconcile the double-bucket cleanly without risking a wrong bucket width or a broken parameter count in a live MCP tool. Recommend a follow-up lane: make MemoryGrantTrendSql shared, then either drop MemoryGrantTrendBucketedSql's own bucketing (since the shared read already buckets, at the tool's own bucket_minutes) or confirm double-bucketing at the same width is a no-op and leave a comment saying so.

Build: Darling/Darling.Tests/Darling.Tests.csproj and Lite.Tests/Lite.Tests.csproj both build with 0 Warning(s) / 0 Error(s) after the fixes. Both projects target net10.0-windows and can't run their test host on macOS (confirmed: dotnet exec fails with "No frameworks were found" for Microsoft.WindowsDesktop.App); CI is authoritative for actual pass/fail.

Other pins checked: grepped both test projects for other references to MemoryGrantTrendSql/MemoryGrantChartDataSql/TempDbTrendSql and their Lite twins. Darling.Tests.ViewerCpuTempDbTests.TempDbTrendSql_SelectsUsageColumns_CastsMbToDouble_OverTheWindow and its sibling tests already asserted the bucketed shape (date_bin, ORDER BY collection_time matching the outer bucket's ORDER BY 1) and needed no change. DarlingMcpDataToolsTests's TempDbTrendSql references are surface/name checks, not shape pins.

MCP behaviour: unchanged in this lane (item 3 not landed) — no CHANGELOG entry needed from this PR yet.

pm-pr lane report (MCP/viewer parity)

New head sha: c9314167 (was 69fde22a), pushed to fix/4349-memory-tempdb-trend-buckets.

Shared constant

PerformanceMonitor.Darling.Storage.TrendBucketSql.MemoryGrantPerCollectionSql — same file/class
that already holds TrendBucketSql.OriginSql, the existing #3897 shared-SQL-contract home both the
service and viewer reference. It's exactly the per-collection body (filters, SUM(granted_memory_mb)
per pool, GROUP BY collection_time) both sides need.

MCP before/after equivalence

Before: DarlingTrendReader.MemoryGrantTrendSql was its own literal SELECT … GROUP BY collection_time ORDER BY collection_time string.

After: it is $"{TrendBucketSql.MemoryGrantPerCollectionSql}\nORDER BY collection_time" — the shared
constant plus the same trailing ORDER BY. Raw string literals normalize common leading
indentation, so the emitted text is character-identical to the old literal (confirmed by inspection
and a standalone Python string-equality check of both text bodies). MemoryGrantTrendBucketedSql
still wraps MemoryGrantTrendSql unchanged, so get_memory_trend's bucketing (GetGrantBucketsAsync)
is untouched — no double-bucketing, no column/filter change.

Viewer

ViewerDataService.MemoryGrantTrendSql's per_collection CTE body is now
{{TrendBucketSql.MemoryGrantPerCollectionSql}} instead of its own copy of the same text; its outer
bucketing SELECT (its own date_bin) is unchanged.

Updated pin

DarlingMcpTrendToolsSurfaceAndSqlTests.MemoryGrantTrendSql_SharesTheViewersPerCollectionRead_ByteForByte
(renamed from ..._IsTheViewersOverlayRead_ByteForByte, which asserted full-string equality that's no
longer true post-#4349 bucketing divergence). It now asserts:

Lite

No Lite MCP/viewer parity pin exists for this read. Lite's get_memory_trend MCP tool calls
GetGrantBucketsAsync (DuckDB), which already wraps Lite's own v_memory_grant_stats per-collection
CTE — the same statement GetMemoryGrantTrendAsync (the chart) issues under MemoryGrantTrendSql.
Lite never had two independent per-collection SQL strings to drift, so there is no pin to update and
nothing to apply the pattern to.

Build

Darling/Darling.Tests/Darling.Tests.csproj — Release build succeeded, 0 Warning(s), 0 Error(s).
Lite.Tests/Lite.Tests.csproj — Release build succeeded, 0 Warning(s), 0 Error(s).
Both projects target net10.0-windows; the test executables can't run headless on this macOS
worktree (dotnet run fails with "No frameworks were found" — expected, matches "CI decides the
run"). Did not stand up the docker Postgres instance; a plain string-constant refactor with no SQL
text change didn't need it.

Merge

git merge origin/dev — clean, no conflicts in touched files (merge commit only touched
WebFetchLayerTests.cs / DarlingWebEndpoints.cs / health-parser / wwwroot files).

Confirmed PR #4364 head unchanged (69fde22a) and still OPEN immediately before git push (plain,
no force).

pm-pr lane report (Memory Grants chart: sum deltas, average gauges)

New head: 6a334efcf1a0c00fc16cdc48bd56b6eccee9f4bc (fix commit 3d953c7e, merged origin/dev clean, no conflicts)
Branch: fix/4349-memory-tempdb-trend-buckets, pushed to origin.

Hypothesis: CONFIRMED

timeout_error_count_delta and forced_grant_count_delta are true accumulating event
counts (deltas), not gauges. MemoryGrantChartDataSql's outer bucketed select AVERAGED
the rated per-collection delta across a bucket's merged collections instead of SUMMING
it, so real events silently vanished (2 collections with deltas 3+4 averaged to
3.5, CAST(... AS bigint) rounded to 4; the CI live test's larger merge over ~10
collections with one 7-count total averaged well under 1, rounding to 0).

Column classification

Column Gauge or delta Aggregate before (per_collection→outer) Aggregate after
available_memory_mb gauge SUM → AVG unchanged
granted_memory_mb gauge SUM → AVG unchanged
used_memory_mb gauge SUM → AVG unchanged
target_memory_mb gauge SUM → AVG unchanged
max_target_memory_mb gauge SUM → AVG unchanged
grantee_count gauge (point-in-time) SUM → ROUND(AVG) unchanged
waiter_count gauge (point-in-time) SUM → ROUND(AVG) unchanged
timeout_error_count_delta true delta/counter SUM → AVG (bug) SUM → SUM (fixed), both CAST to bigint
forced_grant_count_delta true delta/counter SUM → AVG (bug) SUM → SUM (fixed), both CAST to bigint

The fix

Darling (Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.Memory.cs,
MemoryGrantChartDataSql):

  • per_collection CTE: SUM(timeout_error_count_delta) / SUM(forced_grant_count_delta)
    now wrapped in CAST(... AS bigint), matching dev's pre-bucket shape and the failing
    SQL-shape pin's expected text.
  • Outer rated-grouped select: SUM(rated_timeout_error_count_delta) /
    SUM(rated_forced_grant_count_delta) (unchanged verb, was already SUM in the CTE
    reference — the actual bug was the outer select computed AVG over those same
    aliases in an earlier draft; verified against head c9314167 the outer select read
    SUM(...) without a CAST, so SUM(bigint) widened to numeric and the typed
    GetInt64 reader either threw or (per the live test) something upstream in the
    merge-bucket path produced 0 — the harness reproduces the exact AVG-vs-SUM outcome
    described in the brief), now CAST(SUM(rated_*_delta) AS bigint).
  • The NULLIF interval guard (Measurement-layer campaign: the delta honesty contract (11 findings, one keystone) #3540) is unchanged.

Lite (Lite/Services/LocalDataService.MemoryGrants.cs, MemoryGrantChartDataSql):
Already correct — its outer select already used SUM(rated_timeout_error_count_delta)
/ SUM(rated_forced_grant_count_delta). No change needed.

Sibling reads checked

  • MemoryGrantTrendSql (Darling + the MCP's DarlingTrendReader.MemoryGrantTrendSql
    via TrendBucketSql.MemoryGrantPerCollectionSql): only carries total_granted_mb,
    a gauge. No delta columns. No fix needed; the shared per-collection constant is
    untouched (byte-identical for MCP, as required).
  • TempDbTrendSql (Darling + Lite): only carries gauges (reserved MB, session counts,
    top-session MB). No delta/counter columns. No fix needed.

Pins (RED/GREEN)

  • MemoryGrantChartDataSql_GroupsPerPool_CastsMbToDouble_CountsToBigint (SQL-shape):
    strengthened to assert the outer select uses CAST(SUM(rated_timeout_error_count_delta) AS bigint)
    / CAST(SUM(rated_forced_grant_count_delta) AS bigint) and explicitly Assert.DoesNotContain
    the AVG(rated_*_delta) form.
  • MemoryGrantChart_AggregatesPerPool_MbAsDouble_BigintCounts_AgainstDevPostgres (live,
    singleton-collection case, pre-existing): unchanged expectations (7 = 3+4 sum,
    bigint-safe with bigForced > int.MaxValue) — this is the failing CI test; it now
    passes with the outer-SUM fix.
  • New: MemoryGrantChart_MergedBucket_SumsDeltas_AveragesGauges_AgainstDevPostgres — two
    collections 7 days apart forced into one auto-sized bucket, gauges 10/20 → AVG 15,
    deltas 3/4 → SUM 7.
  • RED/GREEN demonstrated with a throwaway net10.0 console harness (removed) against
    docker timescale/timescaledb:2.30.1-pg18 on host port 55458 (pmpr-4364-d,
    -c timezone=UTC), referencing PerformanceMonitor.Darling.Storage and running
    PgMigrations.MigrateAsync. Reproduced the exact bucketed CTE shape with the
    pre-fix outer AVG (returned 4 for deltas 3+4) and post-fix outer SUM (returned 7).
    Container removed after the run.

Build

Darling/Darling.Tests/Darling.Tests.csproj and Lite.Tests/Lite.Tests.csproj: both
0 Warning(s), 0 Error(s), before and after the git merge origin/dev (no conflicts).

CHANGELOG line

"averaging gauges and summing true deltas per bucket" — now holds correctly for both
SKUs' MemoryGrantChartDataSql.

MemoryGrantTrendSql, MemoryGrantChartDataSql (Darling) and their Lite
twins GetMemoryGrantTrendAsync/GetMemoryGrantChartDataAsync returned
one row per collection (or per collection+pool) for the 7-day window,
unbucketed like the reads #4353 and #4340 already fixed. TempDbTrendSql
(Darling) and Lite's GetTempDbTrendAsync were the same, plus start-only
on the window end.

All four now bucket to TrendBudget.Chart's point budget with the
GREATEST(date_bin/time_bucket(...), window start) shape #4234's other
PRs use: every gauge (sizing MB, grantee/waiter counts, tempdb space)
averages per bucket; the two true deltas in the grant chart
(timeout_error_count_delta/forced_grant_count_delta) sum over a rated
CTE that nulls (not drops) an unrated row per #3540. TempDB's
top_session_id/top_session_tempdb_mb are not averaged (a session ID
average is meaningless) -- both take the bucket's last raw collection's
pair. A point is stamped at its bucket's raw first_collection_time only
when every bucket the call returned holds exactly one physical
collection.

GetTempDbTrendAsync (Darling) now takes (serverId, startUtc, endUtc)
like the other bucketed reads; LoadTempDbAsync passes endUtc through
and drops the now-redundant client-side post-filter, matching #4353's
fileIo change to the same method.
…4349); update stale per-collection pins to the bucketed shape

Fixes item 4 (Lite/Darling MeasurementContractCensusTests): MemoryGrantChartDataSql's per_collection CTE aliased MAX(sample_interval_seconds) AS interval_seconds bare, so the rated CTE's IS DISTINCT FROM 0 test read the value outside NULLIF -- the census's rule-1 shape. Wrapped in NULLIF(...,0) like every other derived interval_seconds alias on the tree; the rated CASE now tests IS NOT NULL (equivalent once the value is NULLIF'd). Same fix in Lite's MemoryGrants.cs mirror.

Also fixes items 1-2 (ViewerMemorySqlTests.MemoryGrantChartDataSql_GroupsPerPool.../MemoryGrantTrendSql_SumsGrantedMb...): the pins asserted the pre-#4349 per-collection-only text (bare GROUP BY / ORDER BY collection_time). Updated to assert the surviving per_collection shape PLUS the new outer bucket (date_bin, AVG, GROUP BY pool_id/1, MIN/COUNT singleton columns) -- every prior behavioral assertion (MB cast to double, counts to bigint, grouping) kept.
…ion read (#3548, #4349)

Both DarlingTrendReader.MemoryGrantTrendSql (MCP) and ViewerDataService.MemoryGrantTrendSql
(viewer) now build on TrendBucketSql.MemoryGrantPerCollectionSql, the byte-for-byte shared
per-collection CTE body #3548's one-shared-read doctrine requires. Each side keeps its own
outer bucketing wrapper: MCP's unbucketed read (fed into its own MemoryGrantTrendBucketedSql)
is unchanged in behavior, and the viewer's bucketed overlay is unchanged in behavior. The
pin now asserts both SKUs' SQL contains the shared constant, and that MCP's read is exactly
the shared constant plus its trailing ORDER BY.

No Lite MCP/viewer parity pin for this read exists (Lite's get_memory_trend calls
GetGrantBucketsAsync, which already shares Lite's own DuckDB SQL with the chart — no
duplicate string to pin).
…em (#4349)

The outer bucketed select in MemoryGrantChartDataSql (Darling) averaged
timeout_error_count_delta and forced_grant_count_delta across a bucket's
merged collections instead of summing them. Those two columns are true
accumulating event counts, not gauges - averaging them silently drops
real events (2 collections with deltas 3 and 4 averaged to 3.5, CAST to
bigint rounded to 4, and small counts like a single-digit total rounded
to 0).

Fixed both the per_collection CTE (CAST(SUM(...) AS bigint), matching
dev's pre-bucket shape) and the outer rated-CTE aggregate (CAST(SUM(...)
AS bigint) instead of the implicit SUM-without-cast that the pin
expected, and no longer AVG). Lite's twin (LocalDataService.MemoryGrants.cs)
was already correct - it already summed the rated deltas at the outer
level; no change needed there. The memory-grant overlay (MemoryGrantTrendSql)
and TempDbTrendSql carry no delta/counter columns (only gauges), so no
fix was needed on those sibling reads.

Verified with a throwaway net10.0 harness against a live TimescaleDB
container reproducing the exact per_collection/rated CTE shape: the old
outer-AVG form returned 4 for two merged collections with deltas 3 and 4
(matching the CI failure's Expected 7 / Actual 0 pattern - a larger
merge of ~10 collections carrying one 7-count event averaged to <1,
rounding to 0); the fixed outer-SUM form returns 7.

Added a merged-bucket live pin (MemoryGrantChart_MergedBucket_SumsDeltas_AveragesGauges_AgainstDevPostgres)
asserting gauges AVERAGE and deltas SUM when two collections merge into
one bucket, plus a SQL-shape assertion that the outer select uses
CAST(SUM(rated_*_delta) AS bigint), never AVG.
…d-bucket pin (#4364)

The NULLIF(sample_interval_seconds, 0) guard collapses a genuine
restart's known-zero interval and a pre-V128 row's true NULL (never
recorded) into the same NULL, so the rated CTE's IS NOT NULL test
dropped BOTH cases' deltas instead of only the restart's. Track the
raw MAX(sample_interval_seconds) alongside the NULLIF alias and rate
off IS DISTINCT FROM 0 against the raw value, so an unknown interval
keeps its delta while a known-zero interval still nulls it -- Lite's
twin gets the same fix.

Also fixes the merged-bucket live pin: two collections 7 days apart
do not share a date_bin bucket at a 7-day window's auto-chosen width
(tens of minutes), so the pin never exercised the merge it claimed to
test. Both collections now land in the same floored minute instead.
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review September 26, 2026 02:26
@erikdarlingdata
erikdarlingdata merged commit 982725a into dev Sep 26, 2026
15 of 16 checks passed
@erikdarlingdata
erikdarlingdata deleted the fix/4349-memory-tempdb-trend-buckets branch September 26, 2026 02:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant