Skip to content

analyze_server/compare_analysis move reading guidance off tools/list (#3898) - #4073

Merged
erikdarlingdata merged 3 commits into
devfrom
feature/3898-content-sql-core-a
Sep 23, 2026
Merged

erikdarlingdata merged 3 commits into
devfrom
feature/3898-content-sql-core-a

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Refs #3898.

Why

Darling's full MCP catalog costs about 330 KB of tools/list before the first question, and Lite's costs about
148 KB. Long tool descriptions carry reading guidance an agent needs, so nothing is deleted. Guidance moves off
tools/list into each tool's own reading guide, which get_tool_guide(tools[], topics[]) serves on request.

This PR converts analyze_server and compare_analysis, the two largest tools in the six-tool analysis surface,
on both Darling and Lite.

What changes

Both tools' bodies mirror each other field for field between Darling and Lite. The lane confirmed this by reading
AnalyzeAsync/AnalyzeServer and CompareAnalysis on both products. The original descriptions were already
byte-identical, and the new heads and tails stay that way.

The analyze_server head states these facts, each checked against the tool body on both products:

  • It needs 24+ hours of a server's total collected history, or it returns insufficient_data
    (MinimumDataHours = 24 on both products).
  • The window ends at as_of (default now) and runs hours_back long. The anomaly baseline is a rolling 30-day
    hour-of-day-by-day-of-week baseline.
  • unavailable means the window itself was never observed, so there is no verdict either way.
  • empty means the window was observed and nothing fired: a true all-clear.
  • An unanchored run persists its findings to the store. An as_of run is exploratory and is not saved
    (context.PersistFindings gates _findingStore.InsertFindingsAsync the same way on both products).
  • Confidence is an evidence score, not a probability or a diagnosis.
  • remediation_command is advisory only. It is never auto-executed.

The compare_analysis head states these facts:

  • It compares a comparison window (hours_back long, ending at as_of or now) against an earlier baseline
    window of the same length.
  • baseline_hours_back sets how far back that baseline starts, from the same anchor. It must exceed
    hours_back, or the call is refused.
  • Verdicts are banded by a per-server stored baseline's sigma where one exists, or by ladder position otherwise
    (band_source/band_rules say which).
  • Rows group by physical cause. BAD_ACTOR_<hash> appearances are reported as plan_cache_churn, never as new
    or resolved issues.
  • One window against one window (N=1 vs N=1) never proves causation.
  • unavailable means BOTH windows had zero facts, not that nothing changed. coverage_caveat flags a
    partly-collected side.
  • For a rolling 30-day anomaly baseline, analyze_server is the right tool instead.

D2 parameter cap. compare_analysis's baseline_hours_back parameter description was 211 characters. That
is over the 200-character cap for a converted tool's parameters. It is now trimmed to 139 characters: the
guardrail part, what the parameter is, how it is measured, and its default. The dropped sentence ("The baseline
period will be the same duration as the comparison period.") now rides on compare_analysis's own tail. A new
regression test pins it there so it cannot be silently lost.

Numbers (served characters, head + pointer)

Tool / parameter Before After
analyze_server (Darling and Lite, identical) 4,160 617
compare_analysis (Darling and Lite, identical) 1,647 611
compare_analysis.baseline_hours_back (both) 211 139

Corrections

None. Before writing the heads, the lane read both tool bodies on both products to check every guardrail fact.
Those facts are the 24-hour minimum, the persistence gate, the window and anchor behavior, and the
baseline_hours_back refusal rule. All matched the original prose.

D4 removals

None. Neither tool's original description carried an issue reference or an anecdote.

Re-pointed pins

The coordinator re-pointed two pins in commit f7c6acd:
DarlingMcpToolsTests.CompareAnalysis_Description_SaysWhatWorseMeans_AndWhatItDoesNot and its Lite twin
CompareAnalysisDispersionTests.Description_SaysWhatWorseMeans_AndWhatItDoesNot. Both read the whole unsplit
Description and check 8 tokens. They now check each token where it lives.

  • Head, because a caller needs these before trusting a verdict: band_source, band_rules, plan_cache_churn,
    coverage_caveat, N=1 vs N=1.
  • Tail, which holds the original text: delta_sigma, families, and cannot show that a change CAUSED anything.

D9 traps

  1. Q: analyze_server has under 24 hours of collected history. What does it return?
    Correct: status: "insufficient_data" with a message. Not zero findings, not "empty".
    Prevents it: the head's Needs 24h+ history, else insufficient_data.
  2. Q: analyze_server returns status: "unavailable". Were there no problems on the server?
    Correct: no verdict either way. The window itself was never observed (a dead collector), not a clean
    bill of health.
    Prevents it: the head's unavailable: window unobserved (dead collector), no verdict either way.
  3. Q: Is "empty" the same outcome as "unavailable"?
    Correct: no. "empty" means the window WAS observed and nothing fired: a true all-clear.
    Prevents it: the head's empty: window observed, nothing fired (a true all-clear).
  4. Q: A finding's confidence is 0.85. Is there an 85% chance the finding is correct?
    Correct: no. Confidence is an evidence score built from corroboration, not a probability, and a finding
    is not a diagnosis.
    Prevents it: the head's confidence is evidence, not probability or diagnosis.
  5. Q: Does an analyze_server call with as_of set to yesterday save its findings the same way an
    unanchored call does?
    Correct: no. An as_of run is exploratory. Its findings are complete, but they are not written to the
    store.
    Prevents it: the head's Unanchored runs persist findings to the store; as_of runs are exploratory, not saved.
  6. Q: A finding carries remediation_command. Has it already been run?
    Correct: no. It is advisory only and is never executed automatically.
    Prevents it: the head's remediation_command is advisory, never auto-executed.
  7. Q: compare_analysis is called with baseline_hours_back=2 and hours_back=4. What happens?
    Correct: the call is refused. The baseline period must be the earlier one, so baseline_hours_back must
    exceed hours_back.
    Prevents it: the head's must exceed hours_back, else refused.
  8. Q: compare_analysis returns status: "unavailable". Does that mean nothing changed?
    Correct: no. It means BOTH windows had zero collected facts, so there is nothing to compare at all.
    Prevents it: the head's unavailable = BOTH windows had zero facts.
  9. Q: compare_analysis shows several keys as "worse" at the same hour yesterday. Did today's change
    cause that?
    Correct: it cannot show that. This is one window against one window (N=1 vs N=1), and ordinary
    day-to-day variance alone can produce that shape.
    Prevents it: the head's N=1 vs N=1: never proves causation.
  10. Q: A verdict's band_source is "absolute" rather than "baseline". Does that mean the key has no
    historical context at all?
    Correct: it means the key has no stored per-server baseline. The verdict rests instead on how far up its
    own severity ladder the value climbed.
    Prevents it: the head's sigma where a per-server baseline exists, else ladder position; band_source/band_rules say which.
  11. Q: One side of a compare_analysis verdict was only partly collected. Is the verdict trustworthy as
    stated?
    Correct: no. coverage_caveat is set on every verdict resting on that side. Read the verdict with that
    caveat before drawing a conclusion.
    Prevents it: the head's coverage_caveat flags a partly-collected side.
  12. Q: A BAD_ACTOR_<hash> key appears as "new" in a compare_analysis result. Is that a new problem?
    Correct: no. BAD_ACTOR_<hash> appearances are always reported as plan_cache_churn, never counted as
    new or resolved.
    Prevents it: the head's BAD_ACTOR_<hash> is plan_cache_churn, never new/resolved.
  13. Q: To look at an incident that happened yesterday at 2pm, does a caller widen analyze_server's
    hours_back?
    Correct: no. Set as_of to the incident's end instead. hours_back stays the window's length.
    Prevents it: the shared as_of parameter's For a past incident set as_of to its end; do not widen hours_back. plus the head's "Window ends at as_of (default now), length hours_back."
  14. Q: Does analyze_server's 30-day anomaly baseline stay pinned to today when as_of is set to a past
    date?
    Correct: no. The baseline moves with as_of, so the findings are the ones that past window deserves.
    Prevents it: the head's Window ends at as_of ..., anomaly baseline: rolling 30-day hour-of-day x day-of-week buckets.
  15. Q: For routine "is this normal for a Tuesday at 2pm" anomaly detection, is compare_analysis the right
    tool?
    Correct: no. That is analyze_server's job (a 30-day rolling baseline). compare_analysis is for an
    explicit two-window comparison.
    Prevents it: the head's closing sentence, For 30-day anomaly baselines use analyze_server.

Test plan

  • Darling.Tests and Lite.Tests build with 0 warnings.
  • Targeted: McpToolsListBudgetTests (Darling, 5/5), McpToolGuide* and DarlingMcpToolsTests (Darling,
    48/48, 1 skip needing a live Postgres rig), McpToolGuide*, McpToolsListBudgetTests and
    CompareAnalysisDispersionTests (Lite, 31/31).
  • Red-watch: broke analyze_server's head guardrail sentence, watched McpToolGuideHeadsSqlCoreTests fail,
    restored it.
  • Dropcheck (dropcheck.py origin/dev HEAD analyze_server compare_analysis): all four rows (both tools,
    both products) show missing=0.
  • Full Darling.Tests and Lite.Tests suites, run once after git merge origin/dev. Not run in this pass.
    See Handoff.
  • Plain-English checker on this PR body. Fixed the real hits (semicolons, long sentences, an em dash). The
    CHANGELOG entry is checked separately. See its own file.

Coordinator verification

  • Merged origin/dev into the branch (merge commit 7c8484b), then added commit f7c6acd, which re-points the two pins.
  • dropcheck: missing=0 for both tools on both products.
  • D9: a small model answered all 15 traps from the served text alone, with no misreads. It answered Q6, Q10, Q11 and Q14 with "not stated" and pointed to the guide.
  • Full suites on f7c6acd: Darling.Tests 13,229 total, 0 failed. Lite.Tests 5,234 total, 0 failed. Both builds had 0 warnings.

Handoff

Stopped at the first context-watchdog warning with both tools fully converted and green on both products.
Nothing is left partially converted.

  1. Pins: done by the coordinator in commit f7c6acd. See Re-pointed pins.
  2. Full suites: done by the coordinator after merging origin/dev (see Coordinator verification).
  3. No topics added. McpToolGuideTopics.SqlCore.cs's s_sqlCore array is left empty. No prose repeats
    across these two tools. This was checked directly against both extracted description texts.

🤖 Generated with Claude Code

https://claude.ai/code/session_01M7r1CXmEcBmsrTV34CCLfF

erikdarlingdata and others added 3 commits September 23, 2026 16:34
…3898)

Splits both tools' descriptions into a terse head (guardrail facts: minimum
history, window/as_of, empty vs unavailable semantics, persistence, advisory
remediation for analyze_server; window definition, banding, causation caveat
for compare_analysis) and a tail served only by get_tool_guide, on both
Darling and Lite (the two products' tool bodies mirror each other
field-for-field, so the original prose was already byte-identical and stays
that way in the head and tail). compare_analysis's baseline_hours_back
parameter description is trimmed from 211 to 139 chars to clear the D2 200
char cap; the dropped sentence rides on the tool's own tail.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M7r1CXmEcBmsrTV34CCLfF
DarlingMcpToolsTests and its Lite twin CompareAnalysisDispersionTests read the
whole unsplit Description. Pin the five tokens a caller needs before trusting a
verdict to the served head, and the other three to the tail.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M7r1CXmEcBmsrTV34CCLfF
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review September 23, 2026 20:53
@erikdarlingdata
erikdarlingdata enabled auto-merge (squash) September 23, 2026 20:53
@erikdarlingdata
erikdarlingdata merged commit d95e74e into dev Sep 23, 2026
16 of 18 checks passed
@erikdarlingdata
erikdarlingdata deleted the feature/3898-content-sql-core-a branch September 23, 2026 21:02
erikdarlingdata added a commit that referenced this pull request Sep 23, 2026
…ntries in their sections (#4080)

Adds 42 entries and 42 link refs (#3992, #3995, #3996, #3998, #4001, #4002, #4003, #4007, #4010, #4011, #4013, #4015, #4020, #4022, #4025, #4029, #4030, #4031, #4036, #4038, #4039, #4040, #4044, #4047, #4048, #4049, #4050, #4051, #4055, #4061, #4063, #4064, #4065, #4066, #4067, #4068, #4069, #4070, #4071, #4073, #4074, #4078). Each PR's entry was buffered, and this lands every entry whose PR was merged on origin/dev when it ran.

#3989 left 26 entries under bare 'Changed' and 'Fixed' lines above '### Added'. They move into '### Changed' and '### Fixed', below the new entries, and one blank line stays under [Unreleased].


Claude-Session: https://claude.ai/code/session_01Ua31ugERL5DmhFVRtf6keQ

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant