Skip to content

get_long_query_completions moves reading guidance off tools/list (#3898) - #4122

Merged
erikdarlingdata merged 4 commits into
devfrom
feature/3898-content-longquery
Sep 24, 2026
Merged

erikdarlingdata merged 4 commits into
devfrom
feature/3898-content-longquery

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 24, 2026 •

Copy link
Copy Markdown
Owner

Refs #3898.

Why

get_long_query_completions served the same 1,037-character description on both Darling and Lite. That text is
the full reading guidance a model needs. Without it, a model can misread the page as newest-first instead of
slowest-first. It can misread a null duration as a bad row instead of an attention, or an empty answer as proof
the collector ran. This PR moves that guidance off tools/list and behind get_tool_guide. Nothing is deleted.

What changes

  • get_long_query_completions converts on both products. It is a twin: desc.py confirmed the served text was
    byte-identical before this PR. The generic cross-SKU lockstep test confirms the new heads still match.
  • New head: 573 characters, served 604 with the pointer, under the 620 target.
  • New pin files: Darling/Darling.Tests/McpToolGuideHeads.LongQuery.cs and
    Lite.Tests/McpToolGuideHeads.LongQuery.cs.
  • Budget updated: the get_long_query_completions tool line in DarlingMcpLongQueryTools.txt and
    McpLongQueryTools.txt (1,037 to 604). No parameter line changed: every parameter description was already
    under the 200-char cap, and none were touched (D8).
  • No SQL, payload, or behavior change. Only the Description(...) string moved.

Served characters (before / after)

Tool Darling before Darling after Lite before Lite after
get_long_query_completions 1,037 604 1,037 604

Corrections

None. The original text's claims (page ranking, the empty-answer advice, the opt-in default) all checked out
against the tool body and its status-routing helpers.

D4 removals

None. The description carried no issue references and no anecdotes.

Re-pointed pins

None found. A git grep for this tool's description text (SLOWEST, completions_returned,
duration-RANKED) found nothing in McpDescriptionTruthPinTests, McpZeroIsAMeasurementTests, or
McpPayloadContractCensusTests on either product. The only existing references to this tool, in
McpPageContractTests and RuntimePreconditionMissTests, check tool shape and status routing, not description
text. Neither needed a change.

Status routes (cited from the code)

The head states one status word literally: empty. The table below also carries the two routes that run before
it. The head's phrase "can mean none in the window, or the collector was off" is scoped to empty only. It does not
claim to cover a not_collected or precondition answer.

Status Darling (file:line) Lite (file:line) Condition
not_collected DarlingMcpLongQueryTools.cs:58 McpLongQueryTools.cs:50 rows.Count == 0 AND the server's probed engine cannot run this collector at all.
precondition DarlingMcpLongQueryTools.cs:63 McpLongQueryTools.cs:55 rows.Count == 0, not_collected did not fire, AND the collector's last recorded run named a specific problem (missing capture session, missing extension, or a degraded/non-fatal skip).
empty DarlingMcpLongQueryTools.cs:64 McpLongQueryTools.cs:56 rows.Count == 0 and neither route above fired: either the collector ran clean with no qualifying rows in the window, or it was off: never enabled (no collection_log row at all, the default for an opt-in collector), or switched off since, with a last recorded run that named no problem.
error DarlingMcpLongQueryTools.cs:108 McpLongQueryTools.cs:101 An exception during the read, via McpHelpers.FormatError, on both products.

D9 traps

  • Q: oldest_returned_event_time and newest_returned_event_time are two minutes apart inside a 24-hour
    window. Does that mean the read only covered those two minutes?
    Correct: No. The page is ranked by duration, not time, so the two stamps only describe how old the
    returned slowest rows are. A short, busy spike can hold every one of the window's slowest completions while
    the underlying read still covered the full 24 hours.
    Prevents it: "THE PAGE IS THE window's limit SLOWEST, NOT ITS NEWEST"
  • Q: newest_returned_event_time is well before the window's end (as_of). Does that mean nothing ran in
    the most recent part of the window?
    Correct: No. It means nothing recent enough made the slowest page. Recent completions can exist but rank
    below the limit-th slowest, or run too fast to qualify at all.
    Prevents it: "THE PAGE IS THE window's limit SLOWEST, NOT ITS NEWEST"
  • Q: truncated is true. Does that mean the page is missing some of the window's very slowest completions?
    Correct: No, the opposite. Because the page is duration-ranked, a truncated page still holds the window's
    slowest rows. What got cut ranked below the limit-th slowest, never above it.
    Prevents it: "truncated means the window held more, none slower than the page"
  • Q: completions_returned is 30, the default limit. Does that number alone say whether the window held
    exactly 30 qualifying completions, or more?
    Correct: No, not by itself. completions_returned caps at limit once the window has at least that many.
    truncated: false means that was the whole window. truncated: true means there were more.
    Prevents it: "This is what bounds the page — read truncated to know whether the window held more." (the limit parameter)
  • Q: A row has duration_ms: null while every other field is populated. Is that a bad or incomplete
    capture?
    Correct: No. It is an attention (a client cancel or query timeout), which has no duration by definition
    and is expected to sort after every timed completion.
    Prevents it: "attentions (cancels/timeouts), ranked duration DESC, attentions (no duration) last"
  • Q: A row's result is "Abort". Is that the same kind of event as an attention?
    Correct: No. An Abort result belongs to a completed rpc/batch event whose query was cancelled mid-run and
    still has a duration. An attention is a separate event type with no duration at all. Both relate to
    cancellation, but they are different captures.
    Prevents it: "attentions (cancels/timeouts), ranked duration DESC, attentions (no duration) last"
  • Q: The call returns status: empty. Does that prove the long_query_completions collector is enabled and
    saw nothing this window?
    Correct: No. Empty is also the answer when the collector was off: never turned on (it is opt-in and OFF by
    default), or switched off since. The answer cannot tell "ran clean" apart from "never ran" beyond what it already says.
    Prevents it: "Collector is opt-in, OFF by default: empty can mean none in the window, or the collector was off"
  • Q: A different call, same tool and server, returns status: precondition, not empty. Does the empty
    answer's "enable it in the schedule" advice still apply?
    Correct: Not directly. precondition means the collector does have a recent recorded run. That run named
    a specific problem: its capture session or extension is missing, or it is degraded. The fix it names can
    differ from just flipping the schedule switch. That advice is scoped to the empty status word, not every
    no-data answer.
    Prevents it: "empty can mean none in the window, or the collector was off"
  • Q: limit=100 was requested and the response holds only 12 completions with truncated: false. Does the
    short page mean the read failed partway through?
    Correct: No. limit is a ceiling, not a target. A quiet window can legitimately hold fewer qualifying
    completions than requested, and truncated: false confirms that is the complete count.
    Prevents it: "truncated means the window held more, none slower than the page"
  • Q: The description says completions come from "the opt-in long-query trace." If a server has never
    turned that on, does it answer status: unavailable?
    Correct: No, unavailable is not one of this tool's routes for a missing trace. A never-enabled, opt-in
    collector answers status: empty, naming the specific collector to turn on, not a server-wide unavailable.
    Prevents it: "Collector is opt-in, OFF by default: empty can mean none in the window, or the collector was off"
  • Q: hours_back is 24 (default) and the answer is empty. If it goes to 168 (a week), does that reliably
    turn up a non-empty answer when the collector is off?
    Correct: No. A wider window only reaches rows captured while the collector was on. If it was never
    enabled, every window comes back empty, because nothing was captured at any hours_back.
    Prevents it: "Collector is opt-in, OFF by default: empty can mean none in the window, or the collector was off"
  • Q: The page holds a row with duration_ms: 45000 first, then the next row shows duration_ms: 100. Is
    the ordering broken?
    Correct: No. Every timed row sorts strictly by duration, descending. A big drop between adjacent rows is
    normal once you are past the true outliers. It is not a sign the sort failed.
    Prevents it: "ranked duration DESC, attentions (no duration) last"

Coordinator verification

  • Head facts fixed: 2.
    1. Scope. "empty can mean none in the window, or never enabled" was narrower than the empty route. Take a
      collector switched off yesterday whose last run named no problem. It also answers empty for the last
      day. The precondition route only fires on a recorded problem. The clause now reads "or
      the collector was off".
    2. D9 round 1 misread Q3. The reader answered Yes: a truncated page is missing some of the very slowest. It
      relied on "whether the window held more". The head now says: "truncated means the window held more, none
      slower than the page." It replaces: "completions_returned/truncated say how many you got and whether
      the window held more." completions_returned stays in the tail, and the limit parameter still says to read
      truncated.
  • Head 586 to 573, served 617 to 604, byte-identical on both products. Pins, budgets and the changelog buffer
    entry are updated to match.
  • D9 round 1 (haiku, 13 questions: Q1-Q12 plus Q99, a collector that ran last week and was switched off
    yesterday): 1 misread (Q3). Q6, Q8, Q9, Q10 and Q12 were answered "can't tell". The rest were correct.
  • D9 round 2 on the fixed head asked Q3, Q7 and Q99 again, plus a new Q13. Q13: truncated is true and the
    last row took 2000 ms. Can a cut row have taken 9000 ms? Round 2 found 0 misreads. Q3, Q7 and Q13 were
    correct, and Q99 was answered "can't tell".
  • Targeted tests after the fix: Darling (McpToolGuide*, McpToolsListBudget*, McpDescription*,
    McpZeroIsAMeasurement*, McpPayloadContractCensus*, MeasurementContractCensus*, McpPageContract*,
    *LongQuery*): 272 total, 0 failed, 8 skipped. RuntimePreconditionMiss*: 3 total, 0 failed, 3 skipped.
    Lite, same classes: 117 total, 0 failed.

Test plan

  • Build Darling.Tests and Lite.Tests: both Build succeeded, 0 Warning(s), 0 Error(s).
  • Targeted, both products (McpToolGuide*, McpToolsListBudget*, McpDescription*,
    McpZeroIsAMeasurement*, McpPayloadContractCensus*): Darling 206 total, 0 failed, 4 skipped. Lite 62
    total, 0 failed.
  • Every class that mentions get_long_query_completions: Darling McpPageContractTests plus
    RuntimePreconditionMissTests: 43 total, 0 failed. Lite McpPageContractTests: 22 total, 0 failed.
  • Red-watch: broke the collector-gate sentence in the Darling head, confirmed
    McpToolGuideHeadsLongQueryTests failed (1 failed), restored it, confirmed green again.
  • Dropcheck: get_long_query_completions [Darling] old=1037 new=1634 sentences=6 missing=0. [Lite]
    old=1037 new=1634 sentences=6 missing=0.
  • Full suite once, after git merge origin/dev (merge commit 23615e6): Darling.Tests Total 13335,
    Errors 0, Failed 2, Skipped 636. Both failures are the known non-elevated
    DarlingInstallLocationTests (ThePreLockWritableExtractionCheck_... and TheInstallTreeLock_...) the
    brief calls out. They are unrelated to this change.
  • Full suite once: Lite.Tests Total 5268, Errors 0, Failed 1. The one failure,
    AnalysisPassTokenThreadingTests.TheReadLockWaitIsAbandonableWhileAWriterHoldsIt, is a threading/lock
    test in Lite's analysis pass with no relation to this PR's file. Re-run alone 3 times, it passed every
    time. It only fails under full-suite concurrency. See Handoff.
  • The plain-English checker ran once on this PR body. The real hits are fixed.

Handoff

  • The full-suite AnalysisPassTokenThreadingTests.TheReadLockWaitIsAbandonableWhileAWriterHoldsIt flake
    (passes in isolation, fails under full-suite load) is unrelated to this PR. It fits the wave's flake census
    better than this lane's narrow scope.
  • No unconverted tools remain in this lane's scope: one tool, one twin pair, fully converted.

🤖 Generated with Claude Code

https://claude.ai/code/session_01QQh8LczP1HGLXQYAx4oTZC

erikdarlingdata and others added 4 commits September 23, 2026 22:51
Refs #3898. Converts get_long_query_completions (a twin, identical on
both products) to the head/tail split: tools/list now serves a 617-char
head (down from 1,037) plus a pointer to get_tool_guide, which serves
the full original text unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QQh8LczP1HGLXQYAx4oTZC
…, and empty covers a collector switched off

D9 round 1 read truncated=true as "some of the very slowest are missing". The page is
ranked duration DESC, so a truncated page cut rows no slower than the page; the head
now says so. "empty can mean none in the window, or never enabled" was too narrow: a
collector switched off yesterday also leaves the last 24 hours empty, so the clause
now reads "or the collector was off". Head 573, served 604 on both products.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NUU29PuGg9TUFBsACgZg2K
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review September 24, 2026 03:58
@erikdarlingdata
erikdarlingdata enabled auto-merge (squash) September 24, 2026 03:58
@erikdarlingdata
erikdarlingdata merged commit ce0cffa into dev Sep 24, 2026
17 of 20 checks passed
@erikdarlingdata
erikdarlingdata deleted the feature/3898-content-longquery branch September 24, 2026 04:32
erikdarlingdata added a commit that referenced this pull request Sep 24, 2026
* Add 48 buffered CHANGELOG entries in one splice

Adds 48 entries and 48 link refs (#4057, #4077, #4079, #4081, #4082, #4083, #4084, #4085, #4086, #4087, #4088, #4089, #4090, #4091, #4092, #4093, #4095, #4096, #4099, #4100, #4101, #4103, #4105, #4106, #4107, #4108, #4109, #4110, #4111, #4113, #4114, #4115, #4116, #4117, #4118, #4119, #4120, #4121, #4122, #4123, #4124, #4125, #4126, #4127, #4131, #4133, #4136, #4141). Each PR's entry was buffered, and this lands every entry whose PR was merged on origin/dev when it ran.

Entries found under a bare section line (none unless the old script ran again) move into the matching ### section, below the new entries.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QQh8LczP1HGLXQYAx4oTZC

* Changelog splice: add #4137's entry, which merged before the splice but was missed

#4137 (pg_plan_capture reads csvlog) merged at 05:57Z with a CHANGELOG entry section in its
body. It was neither spliced nor listed as needing no entry. Its entry goes under Fixed, right
after #4136's, and #4136's last sentence now points to it instead of saying plan capture is
unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NUU29PuGg9TUFBsACgZg2K

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant