Skip to content

Narrow audit_config's collection, bound the regression baseline, bucket PG CPU (#4192, #4195, #4193, #4217) - #4206

Merged
erikdarlingdata merged 11 commits into
devfrom
fix/4192-mcp-slow-tools
Sep 25, 2026
Merged

erikdarlingdata merged 11 commits into
devfrom
fix/4192-mcp-slow-tools

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Closes #4192. Closes #4195. Closes #4193. Closes #4217.

Why

Three MCP read tools were slow or oversized. A Mac-side vendor-style perf review measured each one directly:

What changes

#4192: audit_config

Added a narrow CollectConfigAuditFactsAsync next to each collector's existing family methods: PgFactCollector for SQL Server targets, PgTargetFactCollector for PostgreSQL targets, and Lite's DuckDbFactCollector for parity. Each calls only the specific family methods audit_config actually reads. DarlingAnalysisService and Lite's AnalysisService each got a CollectConfigAuditFactsAsync(serverId, serverName, asOfUtc) wrapper: no coverage witness, no anomaly detector, no scorer. Output is unchanged: same fact keys, same recommendation logic.

#4195: get_query_store_regressions

Bounded the baseline to a fixed 7 days ending at the recent window's start (DarlingQueryStoreRegressionReader.BaselineLookbackDays). Passed as a 6th SQL parameter ($6), a plain timestamp, so TimescaleDB keeps plan-time chunk exclusion. The payload reports baseline_start and baseline_end explicitly. Lite's own copy of this tool had the same unbounded shape and is bounded the same way (LocalDataService.BaselineLookbackDays = 7).

#4193: get_pg_cpu_utilization

Added TrendBuckets.PgCpuMaxPoints (500, the widest cap in the family). Each point covers the CPU/ACU pair with peaks, the capacity trio, and six host-memory columns, about 500 bytes per point. DarlingPgCpuUtilizationReader.HistoryBucketedSql and GetBucketedHistoryAsync date_bin-bucket on sample_time, windowed on collection_time. CPU and ACU are averaged per bucket with each bucket's peak kept beside it.

memory_free_bytes and memory_active_bytes take each bucket's worst sample (min free, max active) rather than an average, so a pressure spike is not smoothed away. Routes through TrendBuckets.Resolve like the rest of the converted trend family. Lite has no PostgreSQL-target tools, so there is no Lite counterpart.

#4217: the WPF viewer's third unbounded copy

ViewerDataService.QueryStoreRegressions.cs's GetQueryStoreRegressionsAsync filtered its deduped_baseline CTE only on collection_time < $2 (the window start), with no lower bound. Its own doc comment said so directly. The viewer does not reference DarlingQueryStoreRegressionReader (a different project) or Lite's LocalDataService (a separate app). It keeps its own BaselineLookbackDays = 7 constant, matching the value the other two already keep independently. Added a $5 baseline-start parameter, computed in C# (startUtc.AddDays(-7)) rather than as a SQL interval. DarlingQueryStoreRegressionReader takes it the same way, for the same TimescaleDB chunk-exclusion reason.

Pinned with a source-text assertion for the new collection_time >= $5 bound (fails against the old shape, which had no $5 at all) and a live-Postgres case. Query 400's only baseline snapshot is 8 days before the window start. Before this fix, it joined and showed a 400% CPU regression. After, the INNER JOIN drops it for lacking a baseline inside the 7-day bound.

What the M1b lane did

The M1 lane hit its context budget partway through verification and left five known-red items plus the live-rig measurements outstanding. M1b fixed all five:

  • A stale <see cref="Rounded"/> doc comment.
  • Two McpToolsListBudgetTests/TotalCeilingBytes fixes.
  • A stale PgTargetMcpSurfaceTests string literal.
  • The ResolveEngineAsync call-site-count pins from 3 to 4 for the new CollectConfigAuditFactsAsync entry point.

M1b also bounded Lite's own regression baseline and got both full suites to 0 failed. No Postgres rig was available, so 691 live tests never ran on this branch.

What this lane (M1c) did

  1. Fixed WPF desktop viewer's Query Store Regressions grid still has an unbounded baseline (#4195 twin) #4217 (see above): the viewer's third unbounded baseline copy, with a source pin and a live case.

  2. Stood up the Postgres rig (port 55964) and measured before/after on this branch's own code. The full and narrow audit_config paths both still exist as separate, independently callable entry points on this build. So does the regression reader's baseline-bound parameter. That gives a same-build, same-fixture, same-connection comparison, not a dual-checkout one. Absolute numbers are from a small local fixture on a loopback rig, not the field's production-scale store. They show the shape of the fix (call-count and row-scan reduction), not the field's absolute 7-15 s / 5-9 s numbers. Those came from a pathological read at millions-of-rows scale this fixture does not reproduce.

    • audit_config: DarlingAnalysisService.CollectAndScoreFactsAsync (the full pass, still reachable directly) vs CollectConfigAuditFactsAsync (the narrow pass), 5 calls each, same minimal PG-target config/database-stats fixture. Full pass: 12.2 ms/call. Narrow pass: 6.2 ms/call. About 2x on a nearly-empty local store. The difference is purely call-count (about 30 family reads vs about 5), not the field's dominant pathological-read cost.
    • get_query_store_regressions: seeded 3 queries over a 30-day baseline (1 row/10 min, about 13k rows) plus a 2-hour recent window with one regression. DarlingQueryStoreRegressionReader.GetQueryStoreRegressionsAsync, 5 calls each. Unbounded-simulated (30-day baseline): 43.6 ms/call. Real 7-day bound: 10.6 ms/call. About 4.1x, tracking the about 4.3x row-count reduction. EXPLAIN (ANALYZE, BUFFERS) on the same two bounds: unbounded 588 shared buffer hits / 39.1 ms execution. 7-day: 141 shared buffer hits / 9.8 ms execution. About 4.2x fewer buffers, 4x faster.
    • get_pg_cpu_utilization: seeded 4 hours of 1-minute rows (240) with the V136 memory columns populated. At the tool's true default arguments (hours_back=4, no bucket_minutes): 59,646 bytes (58.2 KB), 119 points. The field number for the pre-get_pg_cpu_utilization never adopted the #3897 trend-bucket contract: 1-minute rows at any window, 618 KB for a day and 4.1 MB for a week #4193 shape at this same window was 103 KB (roughly 240 raw rows at about 430 bytes per row). That is a measured 43% reduction. The new shape is capped regardless of window width: a full week measured 84.5 KB where the old shape measured 4.1 MB. This does not reach the brief's "under 32 KB" target. The auto-bucketer targets TrendBuckets.McpPointBudget (200 points, shared by the whole trend family). A 240-minute default window lands on a 2-minute bucket width (119-120 points) rather than something coarser. Reaching 32 KB needs roughly 64 points. That target was estimated before this exact bucket-width choice was measured. Changing the auto-bucket-width formula is a product decision by the lane that built get_pg_cpu_utilization never adopted the #3897 trend-bucket contract: 1-minute rows at any window, 618 KB for a day and 4.1 MB for a week #4193. This lane is reporting the real, verified number.
    • Added get_pg_cpu_utilization to TrendPayloadBudgetLiveTests' roster (seeded pg_cpu_utilization, default-budget sweep + largest-answer cap-width check, alongside its SQL Server twin). Its own row is the widest in the family by design. It got its own PgCpuDefaultCeilingBytes (100 KB) in that test file. Raising the shared DefaultCeilingBytes (80 KB) weakens what that ceiling guards for the other ten, lighter trend tools. Measured 84,556 bytes / 169 points at a week on that fixture.
  3. Ran this branch's live test classes for the first time. No rig was up on this branch before this lane. Found and fixed one real regression: DarlingQueryStoreRegressionsLiveTests (in DarlingQueryStoreRegressionsTests.cs) still asserted the old "Shorten hours_back" advice text in the no-baseline branch. get_query_store_regressions' baseline is every retained row before the window (no lower bound): 5–9 s per call on a large store, the slowest read on the web Queries tab #4195 correctly removed that line when it bounded the baseline. Shortening hours_back no longer changes how far back a fixed baseline reaches, so the advice no longer applies. The message rewrite never updated this test's pin. Fixed the assertion to match the current, correct message text.

  4. Ran the full Darling.Tests.exe suite once, with the rig up. Dropped and recreated darlingtest first, then ran git merge origin/dev (already up to date). Result: 13,766 total, 0 errors, 3 failed, 47 skipped, 1 not run, 645.9 s. The named flake (ServerListAndSummaryPlanShapeTests.TheShippedReads_...) did not fail. Three failures are flagged for the coordinator:

    • TrendPayloadBudgetLiveTests.EveryDefaultAnswer_...: passed standalone (verified twice, before and after the roster addition). Failed only inside the full run with "Exception while reading from stream" on get_file_io_trend. Looks like a connection/resource hiccup under full-suite parallel load against a small local rig (100 max_connections, default shared_buffers), not a code defect. Not confirmed.
    • PgTargetAuditConfigTests.APostgresTarget_ProjectsItsConfigPgFacts_WithUnitsAndNoSuggestedValue: expected status "review", got "ok".
    • PgTargetAuditConfigTests.AnAuroraTarget_RendersTheCheckpointingKnobsNotApplicable_AndExcludesThemFromTheCheckedCount: expected "not_applicable", got "ok". Both PgTargetAuditConfigTests cases are scoring-outcome mismatches. Neither ran before (no rig on this branch until now, so both were always in the 691 skipped). CollectConfigAuditFactsAsync calls the same three family methods (CollectConfigFactsAsync/CollectMemoryFactsAsync/CollectVacuumFactsAsync) the full pass already called for a PG target. The cause is not confirmed. Do not assume pre-existing without checking.
    • Lite.Tests.exe was not run this lane. Context budget ran out after the Darling.Tests full run and rig teardown.
  5. Stopped the rig and deleted C:\GitHub\worktrees\rig-m1c when finished.

Test plan

  • dotnet build Darling/Darling.Tests/Darling.Tests.csproj: 0 Warning(s), 0 Error(s) (rebuilt after every change in this lane).
  • dotnet build Lite.Tests/Lite.Tests.csproj: 0 Warning(s), 0 Error(s) (M1b's last build. M1c's only Lite-adjacent change was WPF desktop viewer's Query Store Regressions grid still has an unbounded baseline (#4195 twin) #4217's WPF viewer, which does not affect Lite).
  • Live-PostgreSQL rig stood up (port 55964) and measured. See "What this lane did" above for all numbers.
  • get_pg_cpu_utilization added to TrendPayloadBudgetLiveTests' roster and passing standalone.
  • Live classes in every file this branch's diff touches, run against the rig. Found and fixed one stale assertion (DarlingQueryStoreRegressionsLiveTests). ViewerQueriesLivePostgresTests (with the new WPF desktop viewer's Query Store Regressions grid still has an unbounded baseline (#4195 twin) #4217 case) and DarlingAnalysisServiceEngineRoutingTests pass.
  • Full Darling.Tests.exe (no filter) with the rig up: 13,766 total, 0 errors, 3 failed, 47 skipped, 1 not run, 645.9 s. Three failures flagged above.
  • Full Lite.Tests.exe: not run this lane. Context budget ran out.
  • The known flake named in the brief (ServerListAndSummaryPlanShapeTests.TheShippedReads_...) did not fail.

CHANGELOG entry

SECTION: Fixed
ENTRY:

Generated with Claude Code

https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ

erikdarlingdata and others added 7 commits September 24, 2026 23:04
…regressions, get_pg_cpu_utilization (#4192, #4195, #4193)

audit_config now collects only the config/hardware/memory/database-size families it
projects (both SQL Server and PostgreSQL branches, plus Lite parity), instead of the
full collect+detect+score analysis pass. get_query_store_regressions bounds its
baseline to a fixed 7 days ending at the window start instead of every retained row.
get_pg_cpu_utilization adopts the #3897 TrendBuckets contract (bucket_minutes,
averaged + peak fields, a worst-sample memory-pressure pair) instead of returning
every raw 1-minute Performance Insights row.

Known unresolved before this branch is ready:
- DarlingPgCpuUtilizationReader.cs has a stale `<see cref="Rounded"/>` (the helper it
  named was removed in the tool-layer rewrite) - DocCommentHygieneTests fails on it.
- McpToolsListBudgetTests' tools/list byte ceiling is exceeded by roughly 170 bytes -
  needs a description trim on one of the three tools touched here.
- PgTargetMcpSurfaceTests.AuditConfig_ProjectsConfigPgFacts_ForAPostgresTarget failed
  on the last local run and was not yet diagnosed.
- No live-Postgres rig run yet: before/after timing numbers, EXPLAIN (ANALYZE,
  BUFFERS), and the TrendPayloadBudgetLiveTests roster addition for get_pg_cpu_utilization
  are all still outstanding.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…e-regressions baseline too (#4192, #4195, #4193)

Finishes lane M1's draft: fixes the stale cref, the tools/list budget overage, the
audit_config surface-test regression and two entry-point-count pins the narrowed
collect changed. Also found and fixed a real Lite parity gap: get_query_store_regressions'
baseline was still unbounded there, the same defect #4195 fixed in Darling.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
)

ViewerDataService.GetQueryStoreRegressionsAsync's deduped_baseline CTE filtered
only collection_time < window-start, with no lower bound - a third, separate
unbounded copy of the shape #4195 bounded on the MCP tool and Lite. Adds a
viewer-local BaselineLookbackDays = 7 constant (the viewer cannot reference the
MCP reader's or Lite's project) and a $5 baseline-start parameter computed in C#
so TimescaleDB keeps plan-time chunk exclusion.

Pinned with a source-text assertion for the new lower bound and a live-Postgres
case (query 400) whose only baseline snapshot is 8 days back: before this fix it
would have joined and regressed, after it the INNER JOIN drops it for lacking a
baseline inside the bound.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
Seeds pg_cpu_utilization (one row/minute, V136 memory columns populated) and
adds it to the default-budget sweep and the largest-answer cap-width check,
alongside its SQL Server twin get_cpu_utilization. Its own row is the widest
in the trend family by design (~500 bytes/point once bucketed), so it gets its
own PgCpuDefaultCeilingBytes (100 KB) rather than raising the shared
DefaultCeilingBytes (80 KB) and weakening what that ceiling guards for the
other ten, lighter trend tools. Measured 84,556 bytes / 169 points at a
week on this fixture.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…ession class (#4195)

DarlingQueryStoreRegressionsLiveTests seeded the no-baseline branch and still
asserted the old "Shorten hours_back" advice. #4195 (already on this branch)
correctly dropped that line: the baseline is now a fixed BaselineLookbackDays
window, so neither widening nor shortening hours_back changes how far back it
reaches, and M1/M1b's message rewrite in DarlingMcpQueryStoreRegressionTools
never updated this test's pin. Found running this branch's live classes for
the first time in lane M1c, not by a source-text search, since this file
wasn't itself touched by the earlier lanes' diff.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…ss (#4206)

CollectConfigAuditFactsAsync (PgTargetFactCollector) was missing
CollectServerMetadataFactsAsync, so the ServerMajorVersion registry fact
(with is_aurora=1) was absent when CollectConfigFactsAsync tried to stamp
not_applicable on max_wal_size/checkpoint_timeout for Aurora targets.

DarlingAnalysisService.CollectConfigAuditFactsAsync was not calling
_scorer.ScoreAll on the returned facts; AuditConfig maps fact.Severity
to ok/review/warning, so unscored facts all read "ok" regardless of their
values. Config-family scoring is self-contained (pure threshold checks;
no amplifiers that need wait-stats or blocking), so running ScoreAll over
the narrow set is correct and cheap.

Fixes:
- APostgresTarget_ProjectsItsConfigPgFacts_WithUnitsAndNoSuggestedValue:
  shared_buffers at 128 MB now scores 0.4 → "review" (was "ok")
- AnAuroraTarget_RendersTheCheckpointingKnobsNotApplicable_And...:
  max_wal_size on Aurora now gets not_applicable metadata → "not_applicable"
  (was "ok"); shared_buffers also now scores correctly → "review"

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
erikdarlingdata added a commit that referenced this pull request Sep 25, 2026
…corer (#4206)

The parity test extracted from CollectAndScoreFactsAsync up to ComparePeriodsAsync.
CollectConfigAuditFactsAsync sits between those two methods in DarlingAnalysisService.cs.
After #4206 added _scorer.ScoreAll to the narrow pass, the extraction included two scorer
calls instead of one, failing Assert.Equal(1, CountOf(code, Score)).

ReadBody now stops at CollectConfigAuditFactsAsync when present (Darling only),
falling back to ComparePeriodsAsync when not (Lite).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
erikdarlingdata added a commit that referenced this pull request Sep 25, 2026
…corer (#4206)

CollectConfigAuditFactsAsync sits between CollectAndScoreFactsAsync and
ComparePeriodsAsync in DarlingAnalysisService.cs. When #4206 added a
_scorer.ScoreAll(facts) call to the narrow audit-config pass, the test's
ReadBody extracted past it into CollectConfigAuditFactsAsync, counting 2
scorer calls instead of 1. Stop at CollectConfigAuditFactsAsync when
present (Darling only); Lite's AnalysisService.cs has no such method, so
IndexOf returns -1 and the fallback to ComparePeriodsAsync applies.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review September 25, 2026 05:34
@erikdarlingdata
erikdarlingdata enabled auto-merge (squash) September 25, 2026 05:34
@erikdarlingdata erikdarlingdata changed the title DO NOT MERGE (MCP slow tools): narrow audit_config's collection, bound the regression baseline, bucket PG CPU (#4192, #4195, #4193) Narrow audit_config's collection, bound the regression baseline, bucket PG CPU (#4192, #4195, #4193, #4217) Sep 25, 2026
erikdarlingdata and others added 3 commits September 25, 2026 02:34
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
CI measured 171,719 bytes on this branch; the ceiling inherited from dev was
171,637. +82 bytes net after the four PRs' audit_config narrowing, regression
baseline bounding, and PG CPU bucketing changes.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
@erikdarlingdata
erikdarlingdata merged commit 7cefe66 into dev Sep 25, 2026
15 of 16 checks passed
@erikdarlingdata
erikdarlingdata deleted the fix/4192-mcp-slow-tools branch September 25, 2026 07:08
erikdarlingdata added a commit that referenced this pull request Sep 25, 2026
…seline to 171_719

171_637 + 82 (#4206) + 364 (#4254) = 172_083

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
erikdarlingdata added a commit that referenced this pull request Sep 25, 2026
…#3953) (#4208)

* Keep the latest Query Store snapshot per interval as it is written (#3953): V140 and the writer

V140 adds collect.query_store_interval_latest (one row per Regular interval
identity), its per-server coverage row, and a pending-replay table. The Query
Store COPY applies each batch in its own transaction behind a savepoint with a
5 s lock_timeout; an apply fault rolls back to the savepoint, records the batch
for replay and lets raw commit, so raw ingestion never depends on the table.
Coverage creation and the hourly gap check run before the transaction with
bare-parameter bounds, so nothing inside it reads raw except by the batch's
collection_time. The viewer probe gets the V140 sentinel and top arm.

Checkpoint: the readers do not use the table yet.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv

* PLAN_REGRESSION and its drill-down read the interval table where its coverage holds the raw read's snapshots (#3953)

Both regression reads get a table twin that shares everything above the
dedup with the shipped raw SQL (split into shared constants, byte-identical).
The fact decides the source per server and pass (coverage start, pending
batches, raw's chunk floor, the table's floor) and records it on the
analysis context, so the drill-down reads the same source. Any fault in the
decision reads raw.

A live test seeds real regressions through the write path and pins the fact
and drill-down rows identical from both sources, then shows only the table
still reaching a best plan after raw is purged below it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv

* The interval table and its pending rows purge at 15 days (#3953)

DarlingRetention deletes the latest-snapshot table's intervals past
QueryStoreIntervalLatestRetentionDays (15, on first_execution_time) and
pending-replay rows past the same horizon, failure-isolated like every
sibling. It stays out of RawTierCoverage and RetentionPolicies. The batched
DELETE is the whole path until the table's hypertable conversion lands;
drop_chunks joins then, in collection_log's shape.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv

* Regressed queries carry the best plan's last run, and the force bot skips best plans older than 4 days (#3953)

With the interval table the PLAN_REGRESSION window really reaches 14 days,
so a best plan can be two weeks old. Both drill-downs (Darling's two twins
through the shared tail, and Lite's) append best_plan_last_seen; the shared
extractor carries it onto ForcePlanTarget, and ForcePlanBotPolicy blocks a
target whose best plan last ran more than MaxBestPlanAgeDays (4, raw's
retention) before the pass, with the new reason best_plan_stale. That keeps
the would-force journal in the regime it has been scored in. A target with
no age (a pre-#3953 finding) is not gated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv

* PLAN_REGRESSION says how old the faster plan is, and which source it read (#3953)

Both fact reads (Darling's raw and table twins through the shared suffix,
and Lite's) append the best plan's last run. The fact metadata gains
best_plan_age_days and plan_regression_source (1 = interval table, 0 = raw;
Lite is always 0), and the advice states the age: "the faster plan on
record (it last ran 9 days ago)". The 10 CPU-second floor stays absolute.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv

* Pin that nothing inside the Query Store batch transaction reads raw except by the batch's collection_time (#3953)

Review finding 2 as a test: every statement the apply runs inside the raw
COPY's transaction reads query_store_stats only under server_id = $1 AND
collection_time = $2; the two wider reads run before it with bare-parameter
bounds; the conflict target and batch key render from one identity constant.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv

* V140 carries the top-rung pins, and raw Query Store rows have one write path (#3953)

QueryStoreIntervalLatestRungTests takes the "I am the top rung" claims off
V139's test: the dense ladder (Scripts[^1]), the DDL (three engine-plain
tables, one NULLS NOT DISTINCT unique index whose columns are the writer's
conflict target), and the viewer probe's top arm. It also pins the coverage
claim's first guard: no product source spells a write into query_store_stats,
and no other generic-COPY site handles the Query Store collector; the #1912
slice repair is the named exception.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv

* Fix two tests the merge's V140->V142 renumbering broke

CollectionCaveatsRungTests.TheRungIsRegisteredAtTheTopOfADenseLadder still
asserted it was the top rung (RungVersion == SchemaVersion, Same(V141,
Scripts[^1])), which stopped being true once V142 landed above it -
mirroring the same fix CheckpointsTimedRungTests already carries from
when V141 landed on it.

QueryStoreIntervalLatestWriterTests had one hardcoded Scripts.Single(m =>
m.Version == 140) reference (re-running the rung's own SQL after dropping
its index mid-test) that the standard version-bump obligations list
doesn't cover.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3

* Fix ViewerCollectionCaveatsGateTests for the new V142 top rung

The test built its "V141 sentinel absent" argument array with `i < arity
- 1`, which meant "every parameter except the last one" back when V141
was the newest. V142 (#3953) appended its own parameter after
hasCollectionCaveats, so `arity - 1` now names hasQueryStoreIntervalLatest
instead, and the old test left hasCollectionCaveats itself true while
turning off the wrong sentinel -- masked further by MapProbedSchemaVersion
answering 142 outright once the V142 sentinel is true, regardless of
V141's state. Finds hasCollectionCaveats by name via reflection and
requires every sentinel from it up to be false, not just the literal last
parameter.

Also normalizes two files' line endings back to the repo's CRLF
convention (git's clean filter had already silently fixed the committed
blobs on the merge commit; only the on-disk working copy needed it, which
is what a literal-CRLF pin elsewhere caught).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3

* Renumber V142->V143 (conflict with #4216); CHANGELOG mentions MaxBestPlanAgeDays gate (#4208)

PR #4216 merged to dev using V142, so #4208's migration is renumbered to V143. All
references updated: PgMigrations.cs (Migration ctor + const name + doc comment),
StorageVersion.SchemaVersion, ViewerDataService.cs (probe comment + arm comment +
return value), and all pin/rung tests (QueryStoreIntervalLatestRungTests,
QueryStoreIntervalLatestWriterTests, CollectionCaveatsRungTests,
PostmasterStartTimeRungTests, ViewerCollectionCaveatsTests).

The PR body's CHANGELOG entry now includes a second bullet for the force-plan bot
age gate (ForcePlanBotPolicy.MaxBestPlanAgeDays = 4), pinned by
Assert.Equal(4, ForcePlanBotPolicy.MaxBestPlanAgeDays) in ForcePlanBotPolicyTests.

Build: 0 warnings, 0 errors.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ

* Trigger CI: Build did not run after V143 renumber

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ

* Tighten ReadBody end marker to exclude CollectConfigAuditFactsAsync scorer (#4206)

The parity test extracted from CollectAndScoreFactsAsync up to ComparePeriodsAsync.
CollectConfigAuditFactsAsync sits between those two methods in DarlingAnalysisService.cs.
After #4206 added _scorer.ScoreAll to the narrow pass, the extraction included two scorer
calls instead of one, failing Assert.Equal(1, CountOf(code, Score)).

ReadBody now stops at CollectConfigAuditFactsAsync when present (Darling only),
falling back to ComparePeriodsAsync when not (Lite).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ

* Revert "Tighten ReadBody end marker to exclude CollectConfigAuditFactsAsync scorer (#4206)"

This reverts commit fc1055f.

* ci: trigger build on current HEAD

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ

* Fix duplicate collectionCaveatsOrdinal: remove const shadowing the var from dev

Merging dev into this branch left two declarations of `collectionCaveatsOrdinal`
in ViewerCollectionCaveatsGateTests: the `var` added by V142's dev merge and the
`const int 116` from this branch's original test. Remove the const; use the
Array.FindIndex result throughout.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ

* Fix V142 and V143 rung tests after dev merge added V142 between them

V142 (#4196) landed on dev after the fix/3953-qs-interval-latest branch was
cut. Merging dev bumped V143's ProbeOrdinal from 117 to 118 and made V142 no
longer the top rung.

IndexObjectStatsServerTimeIndexRungTests (V142):
- Remove "I am the top rung" assertions (Assert.Equal(RungVersion, SchemaVersion)
  and Assert.Same(V142, Scripts[^1])). Replace with Assert.True(RungVersion < SchemaVersion).
- Change Assert.Equal(ProbeOrdinal, arity - 1) to Assert.True(ProbeOrdinal < arity - 1).
- Add nextArm check that V143's arm sits above V142's in MapProbedSchemaVersion.
- Update class summary and comments to say claims moved to QueryStoreIntervalLatestRungTests.

QueryStoreIntervalLatestRungTests (V143):
- Bump ProbeOrdinal from 117 to 118 (V142 is now at 117).
- Bump PreviousVersion from 141 to 142 (V142 is the rung directly below V143).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant