Repository navigation
analyze_server/compare_analysis move reading guidance off tools/list (#3898) - #4073
Merged
Merged
Conversation
…3898) Splits both tools' descriptions into a terse head (guardrail facts: minimum history, window/as_of, empty vs unavailable semantics, persistence, advisory remediation for analyze_server; window definition, banding, causation caveat for compare_analysis) and a tail served only by get_tool_guide, on both Darling and Lite (the two products' tool bodies mirror each other field-for-field, so the original prose was already byte-identical and stays that way in the head and tail). compare_analysis's baseline_hours_back parameter description is trimmed from 211 to 139 chars to clear the D2 200 char cap; the dropped sentence rides on the tool's own tail. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M7r1CXmEcBmsrTV34CCLfF
DarlingMcpToolsTests and its Lite twin CompareAnalysisDispersionTests read the whole unsplit Description. Pin the five tokens a caller needs before trusting a verdict to the served head, and the other three to the tail. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M7r1CXmEcBmsrTV34CCLfF
erikdarlingdata
marked this pull request as ready for review
September 23, 2026 20:53
erikdarlingdata
enabled auto-merge (squash)
September 23, 2026 20:53
erikdarlingdata
added a commit
that referenced
this pull request
Sep 23, 2026
…ntries in their sections (#4080) Adds 42 entries and 42 link refs (#3992, #3995, #3996, #3998, #4001, #4002, #4003, #4007, #4010, #4011, #4013, #4015, #4020, #4022, #4025, #4029, #4030, #4031, #4036, #4038, #4039, #4040, #4044, #4047, #4048, #4049, #4050, #4051, #4055, #4061, #4063, #4064, #4065, #4066, #4067, #4068, #4069, #4070, #4071, #4073, #4074, #4078). Each PR's entry was buffered, and this lands every entry whose PR was merged on origin/dev when it ran. #3989 left 26 entries under bare 'Changed' and 'Fixed' lines above '### Added'. They move into '### Changed' and '### Fixed', below the new entries, and one blank line stays under [Unreleased]. Claude-Session: https://claude.ai/code/session_01Ua31ugERL5DmhFVRtf6keQ Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #3898.
Why
Darling's full MCP catalog costs about 330 KB of
tools/listbefore the first question, and Lite's costs about148 KB. Long tool descriptions carry reading guidance an agent needs, so nothing is deleted. Guidance moves off
tools/listinto each tool's own reading guide, whichget_tool_guide(tools[], topics[])serves on request.This PR converts
analyze_serverandcompare_analysis, the two largest tools in the six-tool analysis surface,on both Darling and Lite.
What changes
Both tools' bodies mirror each other field for field between Darling and Lite. The lane confirmed this by reading
AnalyzeAsync/AnalyzeServerandCompareAnalysison both products. The original descriptions were alreadybyte-identical, and the new heads and tails stay that way.
The
analyze_serverhead states these facts, each checked against the tool body on both products:insufficient_data(
MinimumDataHours = 24on both products).as_of(default now) and runshours_backlong. The anomaly baseline is a rolling 30-dayhour-of-day-by-day-of-week baseline.
unavailablemeans the window itself was never observed, so there is no verdict either way.emptymeans the window was observed and nothing fired: a true all-clear.as_ofrun is exploratory and is not saved(
context.PersistFindingsgates_findingStore.InsertFindingsAsyncthe same way on both products).remediation_commandis advisory only. It is never auto-executed.The
compare_analysishead states these facts:hours_backlong, ending atas_ofor now) against an earlier baselinewindow of the same length.
baseline_hours_backsets how far back that baseline starts, from the same anchor. It must exceedhours_back, or the call is refused.(
band_source/band_rulessay which).BAD_ACTOR_<hash>appearances are reported asplan_cache_churn, never as newor resolved issues.
unavailablemeans BOTH windows had zero facts, not that nothing changed.coverage_caveatflags apartly-collected side.
analyze_serveris the right tool instead.D2 parameter cap.
compare_analysis'sbaseline_hours_backparameter description was 211 characters. Thatis over the 200-character cap for a converted tool's parameters. It is now trimmed to 139 characters: the
guardrail part, what the parameter is, how it is measured, and its default. The dropped sentence ("The baseline
period will be the same duration as the comparison period.") now rides on
compare_analysis's own tail. A newregression test pins it there so it cannot be silently lost.
Numbers (served characters, head + pointer)
analyze_server(Darling and Lite, identical)compare_analysis(Darling and Lite, identical)compare_analysis.baseline_hours_back(both)Corrections
None. Before writing the heads, the lane read both tool bodies on both products to check every guardrail fact.
Those facts are the 24-hour minimum, the persistence gate, the window and anchor behavior, and the
baseline_hours_backrefusal rule. All matched the original prose.D4 removals
None. Neither tool's original description carried an issue reference or an anecdote.
Re-pointed pins
The coordinator re-pointed two pins in commit f7c6acd:
DarlingMcpToolsTests.CompareAnalysis_Description_SaysWhatWorseMeans_AndWhatItDoesNotand its Lite twinCompareAnalysisDispersionTests.Description_SaysWhatWorseMeans_AndWhatItDoesNot. Both read the whole unsplitDescriptionand check 8 tokens. They now check each token where it lives.band_source,band_rules,plan_cache_churn,coverage_caveat,N=1 vs N=1.delta_sigma,families, andcannot show that a change CAUSED anything.D9 traps
analyze_serverhas under 24 hours of collected history. What does it return?Correct:
status: "insufficient_data"with a message. Not zero findings, not"empty".Prevents it: the head's
Needs 24h+ history, else insufficient_data.analyze_serverreturnsstatus: "unavailable". Were there no problems on the server?Correct: no verdict either way. The window itself was never observed (a dead collector), not a clean
bill of health.
Prevents it: the head's
unavailable: window unobserved (dead collector), no verdict either way."empty"the same outcome as"unavailable"?Correct: no.
"empty"means the window WAS observed and nothing fired: a true all-clear.Prevents it: the head's
empty: window observed, nothing fired (a true all-clear).Correct: no. Confidence is an evidence score built from corroboration, not a probability, and a finding
is not a diagnosis.
Prevents it: the head's
confidence is evidence, not probability or diagnosis.analyze_servercall withas_ofset to yesterday save its findings the same way anunanchored call does?
Correct: no. An
as_ofrun is exploratory. Its findings are complete, but they are not written to thestore.
Prevents it: the head's
Unanchored runs persist findings to the store; as_of runs are exploratory, not saved.remediation_command. Has it already been run?Correct: no. It is advisory only and is never executed automatically.
Prevents it: the head's
remediation_command is advisory, never auto-executed.compare_analysisis called withbaseline_hours_back=2andhours_back=4. What happens?Correct: the call is refused. The baseline period must be the earlier one, so
baseline_hours_backmustexceed
hours_back.Prevents it: the head's
must exceed hours_back, else refused.compare_analysisreturnsstatus: "unavailable". Does that mean nothing changed?Correct: no. It means BOTH windows had zero collected facts, so there is nothing to compare at all.
Prevents it: the head's
unavailable = BOTH windows had zero facts.compare_analysisshows several keys as "worse" at the same hour yesterday. Did today's changecause that?
Correct: it cannot show that. This is one window against one window (N=1 vs N=1), and ordinary
day-to-day variance alone can produce that shape.
Prevents it: the head's
N=1 vs N=1: never proves causation.band_sourceis"absolute"rather than"baseline". Does that mean the key has nohistorical context at all?
Correct: it means the key has no stored per-server baseline. The verdict rests instead on how far up its
own severity ladder the value climbed.
Prevents it: the head's
sigma where a per-server baseline exists, else ladder position; band_source/band_rules say which.compare_analysisverdict was only partly collected. Is the verdict trustworthy asstated?
Correct: no.
coverage_caveatis set on every verdict resting on that side. Read the verdict with thatcaveat before drawing a conclusion.
Prevents it: the head's
coverage_caveat flags a partly-collected side.BAD_ACTOR_<hash>key appears as "new" in acompare_analysisresult. Is that a new problem?Correct: no.
BAD_ACTOR_<hash>appearances are always reported asplan_cache_churn, never counted asnew or resolved.
Prevents it: the head's
BAD_ACTOR_<hash> is plan_cache_churn, never new/resolved.analyze_server'shours_back?Correct: no. Set
as_ofto the incident's end instead.hours_backstays the window's length.Prevents it: the shared
as_ofparameter'sFor a past incident set as_of to its end; do not widen hours_back.plus the head's "Window ends at as_of (default now), length hours_back."analyze_server's 30-day anomaly baseline stay pinned to today whenas_ofis set to a pastdate?
Correct: no. The baseline moves with
as_of, so the findings are the ones that past window deserves.Prevents it: the head's
Window ends at as_of ..., anomaly baseline: rolling 30-day hour-of-day x day-of-week buckets.compare_analysisthe righttool?
Correct: no. That is
analyze_server's job (a 30-day rolling baseline).compare_analysisis for anexplicit two-window comparison.
Prevents it: the head's closing sentence,
For 30-day anomaly baselines use analyze_server.Test plan
Darling.TestsandLite.Testsbuild with 0 warnings.McpToolsListBudgetTests(Darling, 5/5),McpToolGuide*andDarlingMcpToolsTests(Darling,48/48, 1 skip needing a live Postgres rig),
McpToolGuide*,McpToolsListBudgetTestsandCompareAnalysisDispersionTests(Lite, 31/31).analyze_server's head guardrail sentence, watchedMcpToolGuideHeadsSqlCoreTestsfail,restored it.
dropcheck.py origin/dev HEAD analyze_server compare_analysis): all four rows (both tools,both products) show
missing=0.Darling.TestsandLite.Testssuites, run once aftergit merge origin/dev. Not run in this pass.See Handoff.
CHANGELOG entry is checked separately. See its own file.
Coordinator verification
Handoff
Stopped at the first context-watchdog warning with both tools fully converted and green on both products.
Nothing is left partially converted.
McpToolGuideTopics.SqlCore.cs'ss_sqlCorearray is left empty. No prose repeatsacross these two tools. This was checked directly against both extracted description texts.
🤖 Generated with Claude Code
https://claude.ai/code/session_01M7r1CXmEcBmsrTV34CCLfF