Repository navigation
MCP tools/list serves short description heads, and get_tool_guide serves the reading guides, with a size ratchet on both editions (#3898) - #4048
Merged
Conversation
…g guides, with a budget ratchet on both SKUs (#3898) Phase 0: McpToolsListBudgetTests (Darling + Lite) measure the serialized tools/list and pin the total, each served description and each parameter description as ceilings that only go down. Seam: one <<GUIDE>> marker split in McpSchemaCompat.WithGeminiCompatibleTools; get_tool_guide(tools[], topics[]) on both SKUs. Pilot: the nine get_health_parser_* tools converted on both SKUs, the shared empty-window ladder moved to the system_health_empty_windows topic. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv
…ADMEs; GuidePointer rename clears CA1720 (#3898) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv
…ore, and a converted one as head plus pointer (#3898) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv
… Lite's specificity (#4048) D9's misreading check found a fresh Haiku model, given only the served head, hedged an ungated tool's empty answer as ambiguous, and never cited severe_errors'/significant_waits' gate or floors beside the sentence that mattered. All nine pilot heads carried the same "read status, source_observed and last_captured_at" sentence regardless of whether the tool gates at all. Each head's empty-answer sentence now names the tool's own outcome: the seven gated tools (significant_waits says "floors") tie it to events_in_window and the gate/floors; the two ungated tools (system_health, memory_node_oom) say a recorded absence is a real result with no gate to speak of. Verified field-by-field against each tool's own JSON projection before wording it. Darling's nine tails also get Lite's specificity: the two products return byte-identical field sets per tool (confirmed by diffing every projection), so Lite's field lists carry over as-is; the one exception, significant_waits, already matched and was left alone. Also corrects Darling's system_health tail, which credited "sp_HealthParser" (the Dashboard's proc) for data its own PARSE ON READ tools never touch. Updates the pins this touches: McpToolGuideTests' per-tool empty-answer assertions (both SKUs), the 600-char head target raised to 620 (significant_waits' floors sentence is now 615), McpToolsListBudget.txt's nine tool lines and each product's total ceiling, and McpZeroIsAMeasurementTests' description check, which pinned the exact "last_captured_at" wording the new heads intentionally no longer spell out (the data envelope and the guide topic still carry it). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv
Owner
Author
|
D9 misreading check, run by the coordinator.
Round 1, at 26b212e: 10 of 14 clean.
Fixed at 66a5949 (lane W2):
Round 2, at 66a5949: traps 1-5 all clean, each for the right reason, with no guide call needed:
D9 is satisfied for this family. Later content PRs each carry their own trap list, and the coordinator runs the same check. |
erikdarlingdata
marked this pull request as ready for review
September 23, 2026 16:18
erikdarlingdata
enabled auto-merge (squash)
September 23, 2026 16:18
5 of 6 tasks
erikdarlingdata
added a commit
that referenced
this pull request
Sep 23, 2026
…ntries in their sections (#4080) Adds 42 entries and 42 link refs (#3992, #3995, #3996, #3998, #4001, #4002, #4003, #4007, #4010, #4011, #4013, #4015, #4020, #4022, #4025, #4029, #4030, #4031, #4036, #4038, #4039, #4040, #4044, #4047, #4048, #4049, #4050, #4051, #4055, #4061, #4063, #4064, #4065, #4066, #4067, #4068, #4069, #4070, #4071, #4073, #4074, #4078). Each PR's entry was buffered, and this lands every entry whose PR was merged on origin/dev when it ran. #3989 left 26 entries under bare 'Changed' and 'Fixed' lines above '### Added'. They move into '### Changed' and '### Fixed', below the new entries, and one blank line stays under [Unreleased]. Claude-Session: https://claude.ai/code/session_01Ua31ugERL5DmhFVRtf6keQ Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Sep 23, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #3898.
Phase 3's seam (D1 + D3) plus Phase 0's ratchet, on both SKUs in lockstep (D6), proven on one pilot family. The issue stays open for the content PRs.
Why
tools/listcosts about 100k tokens of context before the first call. The long descriptions carry reading guidance agents need (#3541), so the fix is where that guidance lives, not whether it exists. This PR builds the seam so content PRs can move guidance off the wire family by family, and a ratchet so the catalog cannot regrow.What changes
The marker and the split. The marker is
<<GUIDE>>, written inside a tool's existingDescriptionliteral (McpToolGuide.Marker). It can't occur in prose, andgrep -rn "<<GUIDE>>"lists every converted tool.McpSchemaCompat.WithGeminiCompatibleToolsinPerformanceMonitor.Common/Mcp/McpSchemaCompat.cs. Both hosts register every tool through it.McpToolGuideCatalog, a DI singleton that the registration itself adds.get_tool_guide(tools[], topics[])is a new tool on both SKUs:DarlingMcpToolGuideTools, and Lite'sMcpToolGuideTools. The shared logic is inMcpToolGuide.Render.unknown_tools/unknown_topics, withavailable_topics. They are never thrown.deferred, not cut off;names_not_read.system_health_empty_windows(McpToolGuideTopics).Pilot: the nine
get_health_parser_*tools, both SKUs. They are a good fit: the same roughly 400-character empty-window sentence was repeated in all 18 descriptions. That sentence became the topic, and every tail includes the topic verbatim, so one call answers it.event_timewindow ending atas_of, newest first;get_health_parser_memory_brokerclaimed a notification of(RESOURCE_MEMPHYSICAL_HIGH/LOW), but the gate only ever returns LOW. The tail now says(RESOURCE_MEMPHYSICAL_LOW; the gate returns no other).Phase 0 ratchet:
McpToolsListBudgetTestson each SKU.ListToolsResult: descriptions and input schemas, usingMcpJsonUtilities.DefaultOptions, UTF-8 bytes. The tools are built through the same registration path, over the tool types the host source registers.McpToolsListBudget.txtnext to each test:get_tool_guideanswer.Head pins (D3). The new
McpToolGuideTestspin each pilot head's guardrail facts: the window, the empty-answer guardrail, the per-tool gate or floor, and the pointer. They also pin that no sentence was lost (each tail carries the topic, and the topic names every rung), plus cross-SKU identity of the heads and ofget_tool_guide.Existing pins, re-pointed deliberately:
McpZeroIsAMeasurementTests.DescriptionOf/ToolDescription(Darling, which reads both SKUs' source) now read the head. Everything they pin is a zero-is-a-measurement guardrail.McpDescriptionTruthPinTests.Description(Lite) now reads the head. The edition claim, the operator cut and thecreate_statementcaveat are all guardrails.McpPayloadContractCensusTests(the window-floor clause):window_truncatedand "not a page cut" must be in the head. The rest of the clause may span head and guide, because its "spelled truncated before Brains-review campaign: deferred structural residue (from #3538 / #3539 / #3540 / #3541) #3653" WIRE CHANGE notice goes to the guide tail per D4.Inventory (D8: no renames, nothing consolidated or removed, +1 tool per SKU).
llms.txt(89-159) and Lite's instructions count.get_tool_guide. The pointer on each converted head and the tool's own description carry it. Darling had only 8 characters of headroom.Also in this PR.
get_tool_guidewas added to both READMEs' tool lists: the root table, and a Darling README bullet.Small in-lane fix. The schema censuses (
McpUnknownArgumentGuardTests, LiteMcpSchemaCompatTests) treated every non-primitive as a DI service. That silently droppedint?/long?/bool?parameters, and the newstring[], from the schemas they check. Both now use the sharedMcpServedSchema.IsServiceParameter.Numbers
get_tool_guideitself (each SKU)All four totals are measured with the same harness. For "before", the dev versions of the two pilot files were restored and the
get_tool_guideregistrations removed; the ratchet then failed as it should, with 329,870 > 328,331 and 149,448 > 147,053, which also proves it red on the old shape. The net figures includeget_tool_guide's own cost (about 930 B per SKU). The seam's saving is small by design, because the content PRs deliver the cut.Test plan
Darling.TestsandLite.Testsbuild with 0 warnings, 0 errors, aftergit merge origin/dev. A--no-incrementalrebuild surfaced one CA1720 on myPointerconst, which is nowGuidePointer.git merge origin/dev: 13,082 total, 0 failed, 617 skipped (live PostgreSQL classes, no rig). The first full run caughtDarlingWebEndpointsTests.ReadEndpoints_AreExactlyTheReadToolCatalog_MinusTheDocumentedExclusions:get_tool_guidehad no/api/readmirror. It is now inDarlingWebEndpoints.ExcludedToolNameswith its reason (it reads no data; its catalog exists only in the MCP host), andExcludedToolNames_AreTheNonReadSurfaceToolsis updated to match.git merge origin/dev: 5,198 total, 0 failed.McpZeroIsAMeasurementLivePostgresTests, and the live halves of the health-parser tests): not run, because there was no rig. Only descriptions changed in those tools; no tool body, SQL or payload did.tools/listfrom a running host: not touched, per lane rules. The ratchet builds the list through the host's registration path, and pins that every unconverted tool is served byte-identical to its[Description].D9: known traps for the misreading check (the coordinator runs it)
get_health_parser_severe_errorsreturned no rows: were there no severe errors?Correct: only severity 19+ errors off the benign connection-reset list are listed. With
statusempty andevents_in_window> 0, events were captured and gated out.Prevents it: the head's "Gated: severity 19 or higher only, benign connection-reset error numbers excluded, so a lower-severity error is never listed."
get_health_parser_*read answeredunavailablewithsource_observed: false: is the server clean?Correct: no evidence either way. The session was never read.
Prevents it: the head's "An empty answer is not a clean bill: read status, source_observed and last_captured_at", and topic rung 4.
get_health_parser_significant_waitsis empty: were there no waits?Correct: only real-session, non-BACKUP waits of at least 500 ms on non-idle types are listed.
get_wait_statshas the totals.Prevents it: the head's "Floors: ... at least 500 ms ...; shorter waits are never listed. get_wait_stats has the instance-wide totals."
memory_node_oomis empty withstatusempty andsource_observed: true: is there a gap in the data?Correct: a healthy measurement. The read is ungated, and an OOM that never happened is simply absent.
Prevents it: the head's "Ungated: every recorded OOM is returned.", and topic rung 3.
get_health_parser_cpu_tasksis empty: was there no CPU pressure?Correct: only WARNING results with at least 10 pending tasks are listed.
Prevents it: the head's "Gated: WARNING-state results with at least 10 pending tasks only."
hours_backto 48.Correct: set
as_ofto the incident's end.Prevents it: the
as_ofparameter description's "For a past incident set as_of to its end; do not widen hours_back.", and the head's "window ending at as_of".get_health_parser_memory_broker: were there HIGH notifications?Correct: only LOW is ever returned.
Prevents it: the head's "Gated: RESOURCE_MEMPHYSICAL_LOW notifications only." (the tail was corrected in this PR).
get_query_store_regressionsshows a null percent: no change?Correct: undefined, because the baseline side is 0.
Prevents it: the description's pinned "undefined_percents" and "null percent never sorts as 0".
Correct: it is unrated, not measured.
Prevents it: the pinned "unrated_points" in the trend descriptions.
get_pvs_statsshowspct_of_database: null: is the PVS 0%?Correct: it was unmeasured, or has no denominator. A measured 0 MB is published as 0.
Prevents it: the pinned "pvs_measured says whether the DMV reported a size".
window_truncated: true: were rows paged away?Correct: retention floored the window. It is not a page cut, and
effective_start/effective_hours_backgive the reach.Prevents it: the
McpHelpers.WindowTruncatedDescriptionclause (now head-pinned).audit_config's MAXDOP advice: is it tailored to Standard edition?Correct: the edition is reported, never consulted.
Prevents it: the pinned "NO check branches on it".
get_cpu_utilizationon SQL Server: are the samples every 15 seconds?Correct: one ring-buffer record per minute. The 15-second cadence is the Azure SQL DB source.
Prevents it: the pinned "one RING_BUFFER_SCHEDULER_MONITOR record per minute" and "samples_in_bucket is the measured count".
create_statement: is that the fix?Correct: it corroborates a statement already measured slow and is never a diagnosis. Every row carries the fixed caveat, including the regression risk for other plans.
Prevents it: the pinned "corroboration for a statement already measured slow, never a diagnosis", "every row carries the fixed caveat" and "regression risk for other plans".
D9 round (2026-09-23)
The coordinator's misreading check gave a fresh Haiku model only the served heads and asked 14 trap questions (the list above). Two defects followed from what it did with them, both now fixed on this branch.
Finding 1 - the identical empty-answer sentence hid the gate. All nine heads said "An empty answer is not a clean bill: read status, source_observed and last_captured_at.", regardless of whether the tool gates at all. On
memory_node_oom(ungated), given an empty answer withstatusempty andsource_observed: true, the model hedged "ambiguous, maybe a gap" - wrong: perMcpToolGuideTopics.SystemHealthEmptyWindowsrungs (1)-(3), status empty is a real result, and only rung (4) (status unavailable,source_observedfalse) is no evidence. Onsevere_errorsandsignificant_waits, the model anchored on that sentence and never cited the gate/floors stated right beside it.The fix - each head's empty-answer sentence now names its own outcome, verified field-by-field against the actual
WitnessStatuscall inEmptyAsync(every field named below is serialized on both products' empty answers, for all nine tools):Note: these sentences no longer name
last_captured_atby itself (the old universal sentence did); it still rides on the data envelope and every rung of the empty answer, and the guide topic still explains it. UpdatedMcpZeroIsAMeasurementTests's description check accordingly (it only pinssource_observednow, the one fact all three variants still state).Finding 2 (the prior lane's own flagged deferral) - Darling's tails were vaguer than Lite's. Fixed here rather than deferred further. I diffed every one of the nine tools' JSON projections line-by-line between
DarlingMcpHealthParserTools.csandLite/Mcp/McpHealthParserTools.cs: the two products return byte-identical field-key sets on all nine tools (severe_errors'database_nameis resolved throughDarlingSystemHealthReader.ResolveDatabaseNameon Darling vs. a direct DTO property on Lite, but the JSON key and meaning match). That let Lite's tail field lists carry over to Darling as-is for 8 of the 9;significant_waits' Darling tail already matched Lite's specificity and was left alone. While rewriting, also corrected a factual error: Darling'ssystem_healthtail credited "sp_HealthParser" (the Dashboard's server-side proc) for data that this PARSE-ON-READ tool never touches; Lite's "captured by sp_server_diagnostics" (the actual SQL Server internal component) is accurate for both and is now what Darling's tail says too.Tail field-verification table (fields named in Darling's rewritten tail, confirmed present in Darling's own JSON projection):
Pins updated:
McpToolGuideTests(both SKUs) - replaced the single universal-sentence assertion with a per-toolEmptyAnswerFactstable (mirroring the existingGateFactspattern), and raised the head-length target from 600 to 620 (significant_waits' floors sentence is now 615 chars, still far under D2's 1,000 absolute cap).McpToolsListBudget.txton both SKUs - all ninetool get_health_parser_*lines grew (311-472 before, to 360-615 after) and both products'TotalCeilingByteswere raised to the measured value (Darling 328,331 -> 329,494; Lite 147,053 -> 148,216), with the reason logged in each test file's change log.McpZeroIsAMeasurementTests.EveryHealthParserTool_PublishesTheSourceWitness_AndClimbsTheSharedLadder- dropped itslast_captured_at-in-description assertion (see above).Red-watch: reverted
severe_errors' new sentence back to the old universal text, rebuilt, and confirmed bothEveryPilotHead_ServesTheWindow_TheEmptyGuardrail_AndItsGateandPilotHeads_AndTheGuideTool_AreIdenticalOnBothSkusfailed as expected; restored the text, rebuilt, and confirmed green again.Test totals for this round:
Darling.Tests: 13,082 total, 0 failed, 617 skipped (live PostgreSQL, no rig).Lite.Tests: 5,198 total, 0 failed, 0 skipped.For the coordinator to double-check
The Darling pilot tails keep Darling's older, vaguer prose...Resolved in the D9 round above - Darling's tails now match Lite's specificity, field-verified against Darling's own projections.get_collection_health,get_ag_health, the blocking/deadlockdedup_key,get_sweep_reports,get_query_store_clutter). None of it concerns the pilot. Thededup_keydisplay-name scoping spans three tools and is the next topic candidate, for the blocking family's content PR.deprecated/Dashboardwas not changed. It sharesWithGeminiCompatibleTools, so it now registers an unused catalog; its descriptions have no markers and are served unchanged.🤖 Generated with Claude Code
https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv