Repository navigation
Narrow audit_config's collection, bound the regression baseline, bucket PG CPU (#4192, #4195, #4193, #4217) - #4206
Merged
Merged
Conversation
…regressions, get_pg_cpu_utilization (#4192, #4195, #4193) audit_config now collects only the config/hardware/memory/database-size families it projects (both SQL Server and PostgreSQL branches, plus Lite parity), instead of the full collect+detect+score analysis pass. get_query_store_regressions bounds its baseline to a fixed 7 days ending at the window start instead of every retained row. get_pg_cpu_utilization adopts the #3897 TrendBuckets contract (bucket_minutes, averaged + peak fields, a worst-sample memory-pressure pair) instead of returning every raw 1-minute Performance Insights row. Known unresolved before this branch is ready: - DarlingPgCpuUtilizationReader.cs has a stale `<see cref="Rounded"/>` (the helper it named was removed in the tool-layer rewrite) - DocCommentHygieneTests fails on it. - McpToolsListBudgetTests' tools/list byte ceiling is exceeded by roughly 170 bytes - needs a description trim on one of the three tools touched here. - PgTargetMcpSurfaceTests.AuditConfig_ProjectsConfigPgFacts_ForAPostgresTarget failed on the last local run and was not yet diagnosed. - No live-Postgres rig run yet: before/after timing numbers, EXPLAIN (ANALYZE, BUFFERS), and the TrendPayloadBudgetLiveTests roster addition for get_pg_cpu_utilization are all still outstanding. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…e-regressions baseline too (#4192, #4195, #4193) Finishes lane M1's draft: fixes the stale cref, the tools/list budget overage, the audit_config surface-test regression and two entry-point-count pins the narrowed collect changed. Also found and fixed a real Lite parity gap: get_query_store_regressions' baseline was still unbounded there, the same defect #4195 fixed in Darling. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
) ViewerDataService.GetQueryStoreRegressionsAsync's deduped_baseline CTE filtered only collection_time < window-start, with no lower bound - a third, separate unbounded copy of the shape #4195 bounded on the MCP tool and Lite. Adds a viewer-local BaselineLookbackDays = 7 constant (the viewer cannot reference the MCP reader's or Lite's project) and a $5 baseline-start parameter computed in C# so TimescaleDB keeps plan-time chunk exclusion. Pinned with a source-text assertion for the new lower bound and a live-Postgres case (query 400) whose only baseline snapshot is 8 days back: before this fix it would have joined and regressed, after it the INNER JOIN drops it for lacking a baseline inside the bound. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
Seeds pg_cpu_utilization (one row/minute, V136 memory columns populated) and adds it to the default-budget sweep and the largest-answer cap-width check, alongside its SQL Server twin get_cpu_utilization. Its own row is the widest in the trend family by design (~500 bytes/point once bucketed), so it gets its own PgCpuDefaultCeilingBytes (100 KB) rather than raising the shared DefaultCeilingBytes (80 KB) and weakening what that ceiling guards for the other ten, lighter trend tools. Measured 84,556 bytes / 169 points at a week on this fixture. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…ession class (#4195) DarlingQueryStoreRegressionsLiveTests seeded the no-baseline branch and still asserted the old "Shorten hours_back" advice. #4195 (already on this branch) correctly dropped that line: the baseline is now a fixed BaselineLookbackDays window, so neither widening nor shortening hours_back changes how far back it reaches, and M1/M1b's message rewrite in DarlingMcpQueryStoreRegressionTools never updated this test's pin. Found running this branch's live classes for the first time in lane M1c, not by a source-text search, since this file wasn't itself touched by the earlier lanes' diff. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…ss (#4206) CollectConfigAuditFactsAsync (PgTargetFactCollector) was missing CollectServerMetadataFactsAsync, so the ServerMajorVersion registry fact (with is_aurora=1) was absent when CollectConfigFactsAsync tried to stamp not_applicable on max_wal_size/checkpoint_timeout for Aurora targets. DarlingAnalysisService.CollectConfigAuditFactsAsync was not calling _scorer.ScoreAll on the returned facts; AuditConfig maps fact.Severity to ok/review/warning, so unscored facts all read "ok" regardless of their values. Config-family scoring is self-contained (pure threshold checks; no amplifiers that need wait-stats or blocking), so running ScoreAll over the narrow set is correct and cheap. Fixes: - APostgresTarget_ProjectsItsConfigPgFacts_WithUnitsAndNoSuggestedValue: shared_buffers at 128 MB now scores 0.4 → "review" (was "ok") - AnAuroraTarget_RendersTheCheckpointingKnobsNotApplicable_And...: max_wal_size on Aurora now gets not_applicable metadata → "not_applicable" (was "ok"); shared_buffers also now scores correctly → "review" Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
erikdarlingdata
added a commit
that referenced
this pull request
Sep 25, 2026
…corer (#4206) The parity test extracted from CollectAndScoreFactsAsync up to ComparePeriodsAsync. CollectConfigAuditFactsAsync sits between those two methods in DarlingAnalysisService.cs. After #4206 added _scorer.ScoreAll to the narrow pass, the extraction included two scorer calls instead of one, failing Assert.Equal(1, CountOf(code, Score)). ReadBody now stops at CollectConfigAuditFactsAsync when present (Darling only), falling back to ComparePeriodsAsync when not (Lite). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
…corer (#4206) CollectConfigAuditFactsAsync sits between CollectAndScoreFactsAsync and ComparePeriodsAsync in DarlingAnalysisService.cs. When #4206 added a _scorer.ScoreAll(facts) call to the narrow audit-config pass, the test's ReadBody extracted past it into CollectConfigAuditFactsAsync, counting 2 scorer calls instead of 1. Stop at CollectConfigAuditFactsAsync when present (Darling only); Lite's AnalysisService.cs has no such method, so IndexOf returns -1 and the fallback to ComparePeriodsAsync applies. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
erikdarlingdata
marked this pull request as ready for review
September 25, 2026 05:34
erikdarlingdata
enabled auto-merge (squash)
September 25, 2026 05:34
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
CI measured 171,719 bytes on this branch; the ceiling inherited from dev was 171,637. +82 bytes net after the four PRs' audit_config narrowing, regression baseline bounding, and PG CPU bucketing changes. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
This was referenced Sep 25, 2026
erikdarlingdata
added a commit
that referenced
this pull request
Sep 25, 2026
…seline to 171_719 171_637 + 82 (#4206) + 364 (#4254) = 172_083 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
erikdarlingdata
added a commit
that referenced
this pull request
Sep 25, 2026
…#3953) (#4208) * Keep the latest Query Store snapshot per interval as it is written (#3953): V140 and the writer V140 adds collect.query_store_interval_latest (one row per Regular interval identity), its per-server coverage row, and a pending-replay table. The Query Store COPY applies each batch in its own transaction behind a savepoint with a 5 s lock_timeout; an apply fault rolls back to the savepoint, records the batch for replay and lets raw commit, so raw ingestion never depends on the table. Coverage creation and the hourly gap check run before the transaction with bare-parameter bounds, so nothing inside it reads raw except by the batch's collection_time. The viewer probe gets the V140 sentinel and top arm. Checkpoint: the readers do not use the table yet. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv * PLAN_REGRESSION and its drill-down read the interval table where its coverage holds the raw read's snapshots (#3953) Both regression reads get a table twin that shares everything above the dedup with the shipped raw SQL (split into shared constants, byte-identical). The fact decides the source per server and pass (coverage start, pending batches, raw's chunk floor, the table's floor) and records it on the analysis context, so the drill-down reads the same source. Any fault in the decision reads raw. A live test seeds real regressions through the write path and pins the fact and drill-down rows identical from both sources, then shows only the table still reaching a best plan after raw is purged below it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv * The interval table and its pending rows purge at 15 days (#3953) DarlingRetention deletes the latest-snapshot table's intervals past QueryStoreIntervalLatestRetentionDays (15, on first_execution_time) and pending-replay rows past the same horizon, failure-isolated like every sibling. It stays out of RawTierCoverage and RetentionPolicies. The batched DELETE is the whole path until the table's hypertable conversion lands; drop_chunks joins then, in collection_log's shape. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv * Regressed queries carry the best plan's last run, and the force bot skips best plans older than 4 days (#3953) With the interval table the PLAN_REGRESSION window really reaches 14 days, so a best plan can be two weeks old. Both drill-downs (Darling's two twins through the shared tail, and Lite's) append best_plan_last_seen; the shared extractor carries it onto ForcePlanTarget, and ForcePlanBotPolicy blocks a target whose best plan last ran more than MaxBestPlanAgeDays (4, raw's retention) before the pass, with the new reason best_plan_stale. That keeps the would-force journal in the regime it has been scored in. A target with no age (a pre-#3953 finding) is not gated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv * PLAN_REGRESSION says how old the faster plan is, and which source it read (#3953) Both fact reads (Darling's raw and table twins through the shared suffix, and Lite's) append the best plan's last run. The fact metadata gains best_plan_age_days and plan_regression_source (1 = interval table, 0 = raw; Lite is always 0), and the advice states the age: "the faster plan on record (it last ran 9 days ago)". The 10 CPU-second floor stays absolute. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv * Pin that nothing inside the Query Store batch transaction reads raw except by the batch's collection_time (#3953) Review finding 2 as a test: every statement the apply runs inside the raw COPY's transaction reads query_store_stats only under server_id = $1 AND collection_time = $2; the two wider reads run before it with bare-parameter bounds; the conflict target and batch key render from one identity constant. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv * V140 carries the top-rung pins, and raw Query Store rows have one write path (#3953) QueryStoreIntervalLatestRungTests takes the "I am the top rung" claims off V139's test: the dense ladder (Scripts[^1]), the DDL (three engine-plain tables, one NULLS NOT DISTINCT unique index whose columns are the writer's conflict target), and the viewer probe's top arm. It also pins the coverage claim's first guard: no product source spells a write into query_store_stats, and no other generic-COPY site handles the Query Store collector; the #1912 slice repair is the named exception. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv * Fix two tests the merge's V140->V142 renumbering broke CollectionCaveatsRungTests.TheRungIsRegisteredAtTheTopOfADenseLadder still asserted it was the top rung (RungVersion == SchemaVersion, Same(V141, Scripts[^1])), which stopped being true once V142 landed above it - mirroring the same fix CheckpointsTimedRungTests already carries from when V141 landed on it. QueryStoreIntervalLatestWriterTests had one hardcoded Scripts.Single(m => m.Version == 140) reference (re-running the rung's own SQL after dropping its index mid-test) that the standard version-bump obligations list doesn't cover. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3 * Fix ViewerCollectionCaveatsGateTests for the new V142 top rung The test built its "V141 sentinel absent" argument array with `i < arity - 1`, which meant "every parameter except the last one" back when V141 was the newest. V142 (#3953) appended its own parameter after hasCollectionCaveats, so `arity - 1` now names hasQueryStoreIntervalLatest instead, and the old test left hasCollectionCaveats itself true while turning off the wrong sentinel -- masked further by MapProbedSchemaVersion answering 142 outright once the V142 sentinel is true, regardless of V141's state. Finds hasCollectionCaveats by name via reflection and requires every sentinel from it up to be false, not just the literal last parameter. Also normalizes two files' line endings back to the repo's CRLF convention (git's clean filter had already silently fixed the committed blobs on the merge commit; only the on-disk working copy needed it, which is what a literal-CRLF pin elsewhere caught). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3 * Renumber V142->V143 (conflict with #4216); CHANGELOG mentions MaxBestPlanAgeDays gate (#4208) PR #4216 merged to dev using V142, so #4208's migration is renumbered to V143. All references updated: PgMigrations.cs (Migration ctor + const name + doc comment), StorageVersion.SchemaVersion, ViewerDataService.cs (probe comment + arm comment + return value), and all pin/rung tests (QueryStoreIntervalLatestRungTests, QueryStoreIntervalLatestWriterTests, CollectionCaveatsRungTests, PostmasterStartTimeRungTests, ViewerCollectionCaveatsTests). The PR body's CHANGELOG entry now includes a second bullet for the force-plan bot age gate (ForcePlanBotPolicy.MaxBestPlanAgeDays = 4), pinned by Assert.Equal(4, ForcePlanBotPolicy.MaxBestPlanAgeDays) in ForcePlanBotPolicyTests. Build: 0 warnings, 0 errors. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ * Trigger CI: Build did not run after V143 renumber Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ * Tighten ReadBody end marker to exclude CollectConfigAuditFactsAsync scorer (#4206) The parity test extracted from CollectAndScoreFactsAsync up to ComparePeriodsAsync. CollectConfigAuditFactsAsync sits between those two methods in DarlingAnalysisService.cs. After #4206 added _scorer.ScoreAll to the narrow pass, the extraction included two scorer calls instead of one, failing Assert.Equal(1, CountOf(code, Score)). ReadBody now stops at CollectConfigAuditFactsAsync when present (Darling only), falling back to ComparePeriodsAsync when not (Lite). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ * Revert "Tighten ReadBody end marker to exclude CollectConfigAuditFactsAsync scorer (#4206)" This reverts commit fc1055f. * ci: trigger build on current HEAD Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ * Fix duplicate collectionCaveatsOrdinal: remove const shadowing the var from dev Merging dev into this branch left two declarations of `collectionCaveatsOrdinal` in ViewerCollectionCaveatsGateTests: the `var` added by V142's dev merge and the `const int 116` from this branch's original test. Remove the const; use the Array.FindIndex result throughout. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ * Fix V142 and V143 rung tests after dev merge added V142 between them V142 (#4196) landed on dev after the fix/3953-qs-interval-latest branch was cut. Merging dev bumped V143's ProbeOrdinal from 117 to 118 and made V142 no longer the top rung. IndexObjectStatsServerTimeIndexRungTests (V142): - Remove "I am the top rung" assertions (Assert.Equal(RungVersion, SchemaVersion) and Assert.Same(V142, Scripts[^1])). Replace with Assert.True(RungVersion < SchemaVersion). - Change Assert.Equal(ProbeOrdinal, arity - 1) to Assert.True(ProbeOrdinal < arity - 1). - Add nextArm check that V143's arm sits above V142's in MapProbedSchemaVersion. - Update class summary and comments to say claims moved to QueryStoreIntervalLatestRungTests. QueryStoreIntervalLatestRungTests (V143): - Bump ProbeOrdinal from 117 to 118 (V142 is now at 117). - Bump PreviousVersion from 141 to 142 (V142 is the rung directly below V143). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #4192. Closes #4195. Closes #4193. Closes #4217.
Why
Three MCP read tools were slow or oversized. A Mac-side vendor-style perf review measured each one directly:
audit_configanswers from 8 point-in-time facts (CTFP, MAXDOP, max memory, max worker threads, edition, hardware, physical memory, database size). It ran the full analysis engine to get them. That engine collects roughly 30 fact families, including the plan-regression read measured at 4.8 s mean / 6.5 s max on a production store. It also runs the anomaly detector's baseline reads and the scorer. Measured at 7-15 s per call on the web Config tab.get_query_store_regressionscompared a recent window against a baseline defined as everyquery_store_statsrow the store still retains before the window, with no lower bound. The read's cost tracked retention, not the window asked for. Measured at 5-9 s per call, about 70k shared blocks read per call, on a fleet store with about 113M rows.get_pg_cpu_utilizationnever adopted the Trend tools return every raw point: get_file_io_trend sends 12,451 points / 1.4 MB for one server at defaults, more than an LLM client's context #3897 TrendBuckets contract the rest of the trend family (including its SQL Server twin,get_cpu_utilization) already carries. It returned one row per raw 1-minute Performance Insights sample for the whole window regardless of size. That gave 618 KB for a day. With the V136 host-memory columns doubling the row width, a week reached 4.1 MB.ViewerDataService.QueryStoreRegressions.cs) had the same unbounded-baseline shape as the MCP tool before get_query_store_regressions' baseline is every retained row before the window (no lower bound): 5–9 s per call on a large store, the slowest read on the web Queries tab #4195. It was a third, separate copy of the defect, in a different subsystem.What changes
#4192: audit_config
Added a narrow
CollectConfigAuditFactsAsyncnext to each collector's existing family methods:PgFactCollectorfor SQL Server targets,PgTargetFactCollectorfor PostgreSQL targets, and Lite'sDuckDbFactCollectorfor parity. Each calls only the specific family methodsaudit_configactually reads.DarlingAnalysisServiceand Lite'sAnalysisServiceeach got aCollectConfigAuditFactsAsync(serverId, serverName, asOfUtc)wrapper: no coverage witness, no anomaly detector, no scorer. Output is unchanged: same fact keys, same recommendation logic.#4195: get_query_store_regressions
Bounded the baseline to a fixed 7 days ending at the recent window's start (
DarlingQueryStoreRegressionReader.BaselineLookbackDays). Passed as a 6th SQL parameter ($6), a plain timestamp, so TimescaleDB keeps plan-time chunk exclusion. The payload reportsbaseline_startandbaseline_endexplicitly. Lite's own copy of this tool had the same unbounded shape and is bounded the same way (LocalDataService.BaselineLookbackDays = 7).#4193: get_pg_cpu_utilization
Added
TrendBuckets.PgCpuMaxPoints(500, the widest cap in the family). Each point covers the CPU/ACU pair with peaks, the capacity trio, and six host-memory columns, about 500 bytes per point.DarlingPgCpuUtilizationReader.HistoryBucketedSqlandGetBucketedHistoryAsyncdate_bin-bucket onsample_time, windowed oncollection_time. CPU and ACU are averaged per bucket with each bucket's peak kept beside it.memory_free_bytesandmemory_active_bytestake each bucket's worst sample (min free, max active) rather than an average, so a pressure spike is not smoothed away. Routes throughTrendBuckets.Resolvelike the rest of the converted trend family. Lite has no PostgreSQL-target tools, so there is no Lite counterpart.#4217: the WPF viewer's third unbounded copy
ViewerDataService.QueryStoreRegressions.cs'sGetQueryStoreRegressionsAsyncfiltered itsdeduped_baselineCTE only oncollection_time < $2(the window start), with no lower bound. Its own doc comment said so directly. The viewer does not referenceDarlingQueryStoreRegressionReader(a different project) or Lite'sLocalDataService(a separate app). It keeps its ownBaselineLookbackDays = 7constant, matching the value the other two already keep independently. Added a$5baseline-start parameter, computed in C# (startUtc.AddDays(-7)) rather than as a SQL interval.DarlingQueryStoreRegressionReadertakes it the same way, for the same TimescaleDB chunk-exclusion reason.Pinned with a source-text assertion for the new
collection_time >= $5bound (fails against the old shape, which had no$5at all) and a live-Postgres case. Query 400's only baseline snapshot is 8 days before the window start. Before this fix, it joined and showed a 400% CPU regression. After, the INNER JOIN drops it for lacking a baseline inside the 7-day bound.What the M1b lane did
The M1 lane hit its context budget partway through verification and left five known-red items plus the live-rig measurements outstanding. M1b fixed all five:
<see cref="Rounded"/>doc comment.McpToolsListBudgetTests/TotalCeilingBytesfixes.PgTargetMcpSurfaceTestsstring literal.ResolveEngineAsynccall-site-count pins from 3 to 4 for the newCollectConfigAuditFactsAsyncentry point.M1b also bounded Lite's own regression baseline and got both full suites to 0 failed. No Postgres rig was available, so 691 live tests never ran on this branch.
What this lane (M1c) did
Fixed WPF desktop viewer's Query Store Regressions grid still has an unbounded baseline (#4195 twin) #4217 (see above): the viewer's third unbounded baseline copy, with a source pin and a live case.
Stood up the Postgres rig (port 55964) and measured before/after on this branch's own code. The full and narrow
audit_configpaths both still exist as separate, independently callable entry points on this build. So does the regression reader's baseline-bound parameter. That gives a same-build, same-fixture, same-connection comparison, not a dual-checkout one. Absolute numbers are from a small local fixture on a loopback rig, not the field's production-scale store. They show the shape of the fix (call-count and row-scan reduction), not the field's absolute 7-15 s / 5-9 s numbers. Those came from a pathological read at millions-of-rows scale this fixture does not reproduce.audit_config:DarlingAnalysisService.CollectAndScoreFactsAsync(the full pass, still reachable directly) vsCollectConfigAuditFactsAsync(the narrow pass), 5 calls each, same minimal PG-target config/database-stats fixture. Full pass: 12.2 ms/call. Narrow pass: 6.2 ms/call. About 2x on a nearly-empty local store. The difference is purely call-count (about 30 family reads vs about 5), not the field's dominant pathological-read cost.get_query_store_regressions: seeded 3 queries over a 30-day baseline (1 row/10 min, about 13k rows) plus a 2-hour recent window with one regression.DarlingQueryStoreRegressionReader.GetQueryStoreRegressionsAsync, 5 calls each. Unbounded-simulated (30-day baseline): 43.6 ms/call. Real 7-day bound: 10.6 ms/call. About 4.1x, tracking the about 4.3x row-count reduction.EXPLAIN (ANALYZE, BUFFERS)on the same two bounds: unbounded 588 shared buffer hits / 39.1 ms execution. 7-day: 141 shared buffer hits / 9.8 ms execution. About 4.2x fewer buffers, 4x faster.get_pg_cpu_utilization: seeded 4 hours of 1-minute rows (240) with the V136 memory columns populated. At the tool's true default arguments (hours_back=4, nobucket_minutes): 59,646 bytes (58.2 KB), 119 points. The field number for the pre-get_pg_cpu_utilization never adopted the #3897 trend-bucket contract: 1-minute rows at any window, 618 KB for a day and 4.1 MB for a week #4193 shape at this same window was 103 KB (roughly 240 raw rows at about 430 bytes per row). That is a measured 43% reduction. The new shape is capped regardless of window width: a full week measured 84.5 KB where the old shape measured 4.1 MB. This does not reach the brief's "under 32 KB" target. The auto-bucketer targetsTrendBuckets.McpPointBudget(200 points, shared by the whole trend family). A 240-minute default window lands on a 2-minute bucket width (119-120 points) rather than something coarser. Reaching 32 KB needs roughly 64 points. That target was estimated before this exact bucket-width choice was measured. Changing the auto-bucket-width formula is a product decision by the lane that built get_pg_cpu_utilization never adopted the #3897 trend-bucket contract: 1-minute rows at any window, 618 KB for a day and 4.1 MB for a week #4193. This lane is reporting the real, verified number.get_pg_cpu_utilizationtoTrendPayloadBudgetLiveTests' roster (seededpg_cpu_utilization, default-budget sweep + largest-answer cap-width check, alongside its SQL Server twin). Its own row is the widest in the family by design. It got its ownPgCpuDefaultCeilingBytes(100 KB) in that test file. Raising the sharedDefaultCeilingBytes(80 KB) weakens what that ceiling guards for the other ten, lighter trend tools. Measured 84,556 bytes / 169 points at a week on that fixture.Ran this branch's live test classes for the first time. No rig was up on this branch before this lane. Found and fixed one real regression:
DarlingQueryStoreRegressionsLiveTests(inDarlingQueryStoreRegressionsTests.cs) still asserted the old "Shorten hours_back" advice text in the no-baseline branch. get_query_store_regressions' baseline is every retained row before the window (no lower bound): 5–9 s per call on a large store, the slowest read on the web Queries tab #4195 correctly removed that line when it bounded the baseline. Shorteninghours_backno longer changes how far back a fixed baseline reaches, so the advice no longer applies. The message rewrite never updated this test's pin. Fixed the assertion to match the current, correct message text.Ran the full
Darling.Tests.exesuite once, with the rig up. Dropped and recreateddarlingtestfirst, then rangit merge origin/dev(already up to date). Result: 13,766 total, 0 errors, 3 failed, 47 skipped, 1 not run, 645.9 s. The named flake (ServerListAndSummaryPlanShapeTests.TheShippedReads_...) did not fail. Three failures are flagged for the coordinator:TrendPayloadBudgetLiveTests.EveryDefaultAnswer_...: passed standalone (verified twice, before and after the roster addition). Failed only inside the full run with "Exception while reading from stream" onget_file_io_trend. Looks like a connection/resource hiccup under full-suite parallel load against a small local rig (100max_connections, defaultshared_buffers), not a code defect. Not confirmed.PgTargetAuditConfigTests.APostgresTarget_ProjectsItsConfigPgFacts_WithUnitsAndNoSuggestedValue: expected status "review", got "ok".PgTargetAuditConfigTests.AnAuroraTarget_RendersTheCheckpointingKnobsNotApplicable_AndExcludesThemFromTheCheckedCount: expected "not_applicable", got "ok". BothPgTargetAuditConfigTestscases are scoring-outcome mismatches. Neither ran before (no rig on this branch until now, so both were always in the 691 skipped).CollectConfigAuditFactsAsynccalls the same three family methods (CollectConfigFactsAsync/CollectMemoryFactsAsync/CollectVacuumFactsAsync) the full pass already called for a PG target. The cause is not confirmed. Do not assume pre-existing without checking.Lite.Tests.exewas not run this lane. Context budget ran out after the Darling.Tests full run and rig teardown.Stopped the rig and deleted
C:\GitHub\worktrees\rig-m1cwhen finished.Test plan
dotnet build Darling/Darling.Tests/Darling.Tests.csproj: 0 Warning(s), 0 Error(s) (rebuilt after every change in this lane).dotnet build Lite.Tests/Lite.Tests.csproj: 0 Warning(s), 0 Error(s) (M1b's last build. M1c's only Lite-adjacent change was WPF desktop viewer's Query Store Regressions grid still has an unbounded baseline (#4195 twin) #4217's WPF viewer, which does not affect Lite).get_pg_cpu_utilizationadded toTrendPayloadBudgetLiveTests' roster and passing standalone.DarlingQueryStoreRegressionsLiveTests).ViewerQueriesLivePostgresTests(with the new WPF desktop viewer's Query Store Regressions grid still has an unbounded baseline (#4195 twin) #4217 case) andDarlingAnalysisServiceEngineRoutingTestspass.Darling.Tests.exe(no filter) with the rig up: 13,766 total, 0 errors, 3 failed, 47 skipped, 1 not run, 645.9 s. Three failures flagged above.Lite.Tests.exe: not run this lane. Context budget ran out.ServerListAndSummaryPlanShapeTests.TheShippedReads_...) did not fail.CHANGELOG entry
SECTION: Fixed
ENTRY:
audit_configno longer runs the full analysis pass ([Narrow audit_config's collection, bound the regression baseline, bucket PG CPU (#4192, #4195, #4193, #4217) #4206]). It read a full ~30-family collection, detector and scorer pass just to answer 8 point-in-time configuration facts, measured at 7-15 seconds per call. It now runs only the specific family reads those 8 facts come from. The facts and recommendations returned are unchanged.get_query_store_regressions' comparison baseline is now a fixed 7 days ([Narrow audit_config's collection, bound the regression baseline, bucket PG CPU (#4192, #4195, #4193, #4217) #4206]). The baseline used to be every Query Store capture the store retained before the requested window. Both cost and the comparison period grew with retention. It is now a fixed 7-day lookback ending at the window's start, reported asbaseline_startandbaseline_endin the response. A regression against something that changed more than 7 days before the window is no longer caught. A store retaining less than 7 days is unaffected. Performance Monitor Lite's own copy of this same tool had the same unbounded baseline and is fixed the same way.get_pg_cpu_utilizationnow returns bucketed points instead of one row per minute ([Narrow audit_config's collection, bound the regression baseline, bucket PG CPU (#4192, #4195, #4193, #4217) #4206]). A wide window (a week, or longer with the V136 host-memory columns) returned several megabytes per call. It now buckets to a point budget the same way the rest of the trend family does. It averages CPU/ACU per bucket while keeping each bucket's peak. The worst (not average) memory-pressure sample per bucket is kept, so a brief spike is not smoothed away.get_query_store_regressionsand is now bounded the same way ([Narrow audit_config's collection, bound the regression baseline, bucket PG CPU (#4192, #4195, #4193, #4217) #4206]). It was a third, separate copy of the defect, in the WPF app rather than the MCP surface, with its own 7-day fixed lookback.REF:
[Narrow audit_config's collection, bound the regression baseline, bucket PG CPU (#4192, #4195, #4193, #4217) #4206]: Narrow audit_config's collection, bound the regression baseline, bucket PG CPU (#4192, #4195, #4193, #4217) #4206
Generated with Claude Code
https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ