Repository navigation
MCP pages now say what bounded them: caps bind to the caller's limit, truncation is detected not inferred, and no page count is called a total (#3541 A3) - #3594
Merged
Conversation
…ncation is observed, no page count is a total (#3541 A3) Six MCP tool groups on both SKUs carried a HIDDEN reader cap -- LIMIT 200 / 50 / 500 under a tool that advertised `limit`, a Take(limit) on top, and the capped count published under a total_* name. An agent that cannot see the code read a 200-row page of a 5,000-event window as the window, and the "widen hours_back" advice could never help because the cap was on rows, not time. get_blocking / get_blocked_process_reports (200), get_deadlocks (50), get_alert_history (+ a hidden dismissed = FALSE filter), get_long_query_completions (200; Darling kept the SLOWEST, Lite kept the NEWEST then re-ranked, so Lite could omit the window's slowest run), get_plan_corrections (200 newest per-cycle re-captures ~ 16h reach regardless of hours_back), get_waiting_tasks (500 / unbounded, bare envelope) and get_wait_stats (50 under a limit up to 1,000). ONE pattern, the one #3287 established for get_collection_log: the cap is the caller's limit bound as a SQL parameter; the tool fetches limit + 1 and reads the extra row as `truncated` rather than inferring it from count >= limit; the page publishes oldest_returned_* / newest_returned_* and names its `order`; the page count is *_returned, never total_*. get_deadlock_detail and get_blocked_process_xml moved their graph/XML predicate into the SQL so limit counts graphs, not rows they would have discarded. get_blocking / get_deadlocks honour #2159's whole-window promise for dedup_key: the fingerprint scan is bounded by a stated 2,000-row ceiling, the payload carries rows_examined / scan_truncated, and a no-match answer on a scan that ran out says so. get_alert_history states and MEASURES its filter: dismissed_excluded and dismissed_excluded_count on every page, include_dismissed to lift it (each row then labelled `dismissed`), and an all-dismissed window names the filter instead of calling itself quiet. Lite's completions read is now duration-ranked in SQL, so both SKUs serve the same population. Descriptions on every touched tool say what bounds the page. Grids keep their caps (the readers default to them). Census tests pin the dialect across both SKUs; Lite's run against a real DuckDB, Darling's live half is gated on DARLING_TEST_PG.
erikdarlingdata
enabled auto-merge (squash)
September 18, 2026 15:36
…lare the new LF-reading pin Two census pins from the first CI run. AsOfWindowAnchorTests requires every LocalDataService call that CAN take asOfUtc to be given it by name, so the anchor's presence is visible to the scan rather than inferred from position; the new GetSlowestLongQueryCompletionsAsync call passed it positionally. RepoFileAdoptionTests holds the exact set of pins that read LF-normalised source; McpPageContractTests is one (its Lite-description anchor spans the line break on get_plan_corrections), so it is declared.
There was a problem hiding this comment.
LGTM — reviewed the core implementation across both SKUs (Darling: DarlingBlockingReader/McpBlockingTools, DarlingAlertReader/McpAlertTools incl. the include_dismissed/dismissed-count logic, DarlingLongQueryReader/McpLongQueryTools, DarlingSessionReader/McpSessionTools, DarlingDataReader/McpDataTools, DarlingPlanCorrectionReader/McpPlanCorrectionTools; Lite: the corresponding LocalDataService..cs and McpTools.cs files) plus DarlingMcpInstructions.cs / McpInstructions.cs.
Specifically verified:
- SQL parameter ordinal binding is correct on every touched query, including the two-branch (server-scoped vs. fleet-wide)
get_alert_historyreads where the$Nposition oflimitand the newinclude_dismissedbool shifts between branches — traced the C# parameter-add order against both SQL consts and it lines up. - The
limit + 1over-fetch /Count > limittruncation pattern is applied consistently and correctly at every touched call site, including the two-stage fingerprint-scan-then-page flow inget_blocking/get_deadlocks/get_deadlock_detail(scan ceiling truncation computed before filtering, page truncation computed after). - The
(dismissed = FALSE OR $N)filter, thedismissed_excluded_countmeasurement query, and theinclude_dismissedplumbing match field-for-field between Darling and Lite. - The new duration-ranked
get_long_query_completionsread (SQL ServerNULLS LAST/ DuckDB equivalent) produces the same population and ordering on both SKUs, fixing the described Lite-only gap. - Lite/Darling parity holds throughout the touched surface; the one remaining asymmetry I checked (Darling's
dedup_keyfingerprint scan has no Lite equivalent onget_blocked_process_reports/get_deadlocks) predates this PR and isn't part of its stated scope. - No literal
LIMITcaps remain in live code paths (only in comments referencing the old defect);OPTION(RECOMPILE)-equivalent concerns don't apply since these are Postgres/DuckDB reads, not T-SQL collectors.
No correctness, security, or parity issues found.
This was referenced Sep 18, 2026
erikdarlingdata
added a commit
that referenced
this pull request
Sep 18, 2026
… counter, a partial name cannot delete the wrong server, an omitted field is never a clear, and a mute that matched nothing says so (#3541 A14) (#3615) * MCP write tools report what happened: every server you add lands in a counter, a partial name cannot delete the wrong server, an omitted field is never a clear, and a mute that matched nothing says so (#3541 A14) add_servers: the per-row status "collides" (#2280) was counted in no summary counter, so a batch with a collided server summarised as failed: 0 and the server went silently unmonitored. Every status now maps to exactly one counter through CounterOfStatus (added / skipped / collided / failed), the envelope carries requested, the four counters sum to it by construction, and the aggregator throws on a status the map does not know. Statuses are constants. remove_server: the read resolver's first-wins partial match deleted whichever sibling sorted first. The write now honors an exact match, or a partial ONLY when unique; an ambiguous name deletes nothing and returns status "ambiguous" with the candidates named (matched_by exact|partial on success). Instructions: the add_servers row still denied the Entra modes #3484 accepted; it now names ServicePrincipal / ManagedIdentity as accepted and the interactive trio as invalid, plus engine / port and the new counters. update_custom_view / update_custom_alert_rule: one write vocabulary for the optional description - omitted = unchanged, empty string = cleared - through a single shared helper (ResolveOptionalText). The view tool used to write an omitted description as NULL; neither tool offered a clear. mute_analysis_finding (both SKUs): the response carries registered and matched_now (stored findings in the mute's scope carrying the hash); status is "muted" or "muted_unmatched", and Darling's swallowed INSERT failure now surfaces as status "error" instead of "muted". The store gains a hash count read (CountStoredFindingsAsync) on both FindingStores; PgFindingStore.MuteStoryAsync returns whether the row landed. A blank hash is refused. Rider (#3594 residual): /api/read/get_alert_history dispatches and advertises include_dismissed. Tests: counters census + sum, unknown-status throw, construction-site literal census, ResolveForRemoval truth table, Entra parser-vs-table pin, description vocabulary pins and cross-tool census, web catalog + dispatch pin, live extensions for the collision batch / ambiguous removal / omitted-vs-cleared description on both update tools / matched_now on Darling, and a new Lite DuckDB class for the mute disclosure with a negative server id. * Lite FindingStore lock-site census: name the seventh read-lock site (CountStoredFindingsAsync, the #3541 A14 matched_now read) TheWritePathsTakeAReadLockOnPurpose_AndSayWhy pins the count of _duckDb.AcquireReadLock( sites so a new lock site is a decision made against the #2455 reason rather than a tidy-up. The new count read takes the READ lock (it is a read; the exclusion it needs is against maintenance only), so the pin moves 6 → 7 with the site named beside it. First CI run's only failure. * CountStoredFindingsAsync states its lock choice at the site (read lock, maintenance-only exclusion, #2455)
erikdarlingdata
added a commit
that referenced
this pull request
Sep 18, 2026
…ors, connections and lock waits land in pg_log_events off the same tail and the same RDS API the deadlock and plan readers use, redacted with the plan parser's own patterns (#3601) The server log is the engine's primary event record and Darling read exactly two shapes out of it. This makes the shared part of those two readers explicit — PgServerLogTail is the one spelling of the pg_read_file tailer both siblings now splice byte-for-byte, PgLogEntryAssembler is the prefix/zone/ companion-line assembly done once — and adds the classifier on top: IPgLogFamilyParser is the seam, three families ship with parsers (error, connection, lock_wait), three are recognised and stored under their own name for #3602/#3603 to structure (temp_file, autovacuum, checkpoint). One cursor, every parser on every line, one identity per entry (raw_line_hash) for the overlapping tail to dedupe on. V129 creates collect.pg_log_events exactly as the generator emits it (one index; the family index is a separate rung for V104's reason). Both transports: PgLogEventsCollector (self-hosted; the body crosses the wire so no family's recogniser is spelled twice) and RdsLogEventIngestor (its own RdsLogSource, marker committed after the write). get_pg_log_events reads it with the #3594/#3613 page contract. Retention 30 d, operator-tunable. Censuses moved: 28 PG collectors, 70 collector tables, 71 hypertables (workflows 73/84), 34 get_pg_* tools, 152 tools.
erikdarlingdata
added a commit
that referenced
this pull request
Sep 18, 2026
…ors, connections and lock waits land in pg_log_events off the same tail and the same RDS API the deadlock and plan readers use, redacted with the plan parser's own patterns (#3601) The server log is the engine's primary event record and Darling read exactly two shapes out of it. This makes the shared part of those two readers explicit — PgServerLogTail is the one spelling of the pg_read_file tailer both siblings now splice byte-for-byte, PgLogEntryAssembler is the prefix/zone/ companion-line assembly done once — and adds the classifier on top: IPgLogFamilyParser is the seam, three families ship with parsers (error, connection, lock_wait), three are recognised and stored under their own name for #3602/#3603 to structure (temp_file, autovacuum, checkpoint). One cursor, every parser on every line, one identity per entry (raw_line_hash) for the overlapping tail to dedupe on. V129 creates collect.pg_log_events exactly as the generator emits it (one index; the family index is a separate rung for V104's reason). Both transports: PgLogEventsCollector (self-hosted; the body crosses the wire so no family's recogniser is spelled twice) and RdsLogEventIngestor (its own RdsLogSource, marker committed after the write). get_pg_log_events reads it with the #3594/#3613 page contract. Retention 30 d, operator-tunable. Censuses moved: 28 PG collectors, 70 collector tables, 71 hypertables (workflows 73/84), 34 get_pg_* tools, 152 tools.
erikdarlingdata
added a commit
that referenced
this pull request
Sep 18, 2026
…ors, connections and lock waits land in pg_log_events off the same tail and the same RDS API the deadlock and plan readers use, redacted with the plan parser's own patterns (#3601) The server log is the engine's primary event record and Darling read exactly two shapes out of it. This makes the shared part of those two readers explicit — PgServerLogTail is the one spelling of the pg_read_file tailer both siblings now splice byte-for-byte, PgLogEntryAssembler is the prefix/zone/ companion-line assembly done once — and adds the classifier on top: IPgLogFamilyParser is the seam, three families ship with parsers (error, connection, lock_wait), three are recognised and stored under their own name for #3602/#3603 to structure (temp_file, autovacuum, checkpoint). One cursor, every parser on every line, one identity per entry (raw_line_hash) for the overlapping tail to dedupe on. V129 creates collect.pg_log_events exactly as the generator emits it (one index; the family index is a separate rung for V104's reason). Both transports: PgLogEventsCollector (self-hosted; the body crosses the wire so no family's recogniser is spelled twice) and RdsLogEventIngestor (its own RdsLogSource, marker committed after the write). get_pg_log_events reads it with the #3594/#3613 page contract. Retention 30 d, operator-tunable. Censuses moved: 28 PG collectors, 70 collector tables, 71 hypertables (workflows 73/84), 34 get_pg_* tools, 152 tools.
Closed
39 tasks done
erikdarlingdata
added a commit
that referenced
this pull request
Sep 18, 2026
…ors, connections and lock waits land in pg_log_events off the same tail and RDS API the deadlock and plan readers use, redacted with the plan parser's own patterns (#3601) (#3646) * PostgreSQL's log is read by a classifier now, not by two regexes: errors, connections and lock waits land in pg_log_events off the same tail and the same RDS API the deadlock and plan readers use, redacted with the plan parser's own patterns (#3601) The server log is the engine's primary event record and Darling read exactly two shapes out of it. This makes the shared part of those two readers explicit — PgServerLogTail is the one spelling of the pg_read_file tailer both siblings now splice byte-for-byte, PgLogEntryAssembler is the prefix/zone/ companion-line assembly done once — and adds the classifier on top: IPgLogFamilyParser is the seam, three families ship with parsers (error, connection, lock_wait), three are recognised and stored under their own name for #3602/#3603 to structure (temp_file, autovacuum, checkpoint). One cursor, every parser on every line, one identity per entry (raw_line_hash) for the overlapping tail to dedupe on. V129 creates collect.pg_log_events exactly as the generator emits it (one index; the family index is a separate rung for V104's reason). Both transports: PgLogEventsCollector (self-hosted; the body crosses the wire so no family's recogniser is spelled twice) and RdsLogEventIngestor (its own RdsLogSource, marker committed after the write). get_pg_log_events reads it with the #3594/#3613 page contract. Retention 30 d, operator-tunable. Censuses moved: 28 PG collectors, 70 collector tables, 71 hypertables (workflows 73/84), 34 get_pg_* tools, 152 tools. * Review + census: the key-tuple redaction runs to the tuple's true close (a value carrying ')' or ')=(' no longer leaks past it), the exclusion shape keeps both key names; the new COPY writer joins the phase/deadline rosters, the reader sets the MCP-read deadline, the README's derived counts move to 28/70/71 (#3601) * Rebase over #3607: the get_pg_* census is 35 with get_pg_logging_audit beside get_pg_log_events (instructions 153/67/thirty-five, runbook 35) * PgLogEventsLivePostgresTests joins the live-postgres collection: it writes the shared store, so it serializes against the other live classes (hygiene pin) * Both live tests tear down through LiveStoreCleanup on their own connection (#1902 ratchet), not on the body's * PgLogFamilies.IsKnown refuses the reserved 'other' word the read's error message already excludes, so family=other is a rejection rather than a silent zero (review note) * Root README edition table and llms.txt say 153 Darling tools (the cross-app inventory pins read them) * Double quotes are not a safe signal: a double-quoted run is kept only after an identifier noun, the ': "…"' / 'at or near "…"' value shapes and 'Failing row contains (…)' go whole, and the pins use the shapes PostgreSQL actually writes (review) * CONTEXT is stored, redacted like detail (V129 gains the column the lock-wait doc already promised), and a double-quoted run after bare whitespace or a newline is evaluated by the allowlist rather than skipped (review)
This was referenced Sep 18, 2026
erikdarlingdata
added a commit
that referenced
this pull request
Sep 18, 2026
… idiom (#3653 Dashboard mirrors) (#3658) - get_blocking / get_deadlocks / get_alert_history: page counts stop posing as totals (*_returned, truncated, limit_applied_by_reader) - #3594's class - Poison Wait: accumulated over the ten-minute window, graded on the shared PoisonWaitEvaluator bars, clear needs an observation - #3593's class - mute_analysis_finding: registered / matched_now / muted_unmatched - #3615's class - CollectorSeverity (0,0) online -> Unknown, '--' - #3635's class - cntr_value_per_second recomputed as a fraction on the read - #3540 A11 - SqlServerAnomalyDetector lab-box markers replaced with what #3616 measured (comment-only)
This was referenced Sep 19, 2026
erikdarlingdata
added a commit
that referenced
this pull request
Sep 19, 2026
…hared helper, the census sweeps the inference class instead of a roster, and four payload rules become facts and inventories (partial #3653) get_pg_deadlocks, get_pg_plan_capture_readiness and get_pg_index_bloat published truncated = rows.Count >= limit, the #3594 inference; the readiness and index tools withheld their summaries on the strength of it for pages that were the whole set, and get_pg_plans beside them Take(limit)'d a read already cut at limit and said nothing. McpHelpers.BoundPage turns (fetched at limit + 1, limit) into (page, truncated) in one place; the four tools are its first consumers and every count downstream is still a count of the page. McpPageContractTests missed the three because EveryPagedTool_ObservesTruncation walks the ten tools #3541 A3 named and its #3653 sibling walks one file - a roster gap, not a regex gap; the matcher catches both spellings and now witnesses the object-initializer one. NoTool_InfersTruncationFromItsCap_OnEitherSku sweeps every tool body on both SKUs with a population floor and fails on the three at dev. McpPayloadContractCensusTests pins the four unpinned rules as far as a census can honestly go: every catch returns through a shared shape (one ad-hoc site rostered) and the two shared shapes - a sentence and a JSON envelope - are an inventory for the vocabulary lane; every bounded parameter reaches a shared validator or one of five inline days_back refusals, and nothing clamps a parameter; every total_* assigned from a Count names a whole set (rostered), every latest-anchored reader projects its anchor (two known exceptions); truncation-key dialects, severity-word spellings and statement-terminal literal LIMITs are exact inventories. The DarlingMcpPgPlanToolsTests pin that handed the builder two rows at limit 2 and expected truncated encoded the inference as its own proof and becomes the boundary pair. Executed here through the built Darling.Tests.dll and against PostgreSQL 18.4 / TimescaleDB 2.28.1; positive controls: the three files at dev state fail the sweep by name, and BoundPage flipped to >= fails the live boundary, the builder pins and the helper theory.
erikdarlingdata
added a commit
that referenced
this pull request
Sep 19, 2026
…hared helper, the census sweeps the inference class instead of a roster, and four payload rules become facts and inventories (partial #3653) (#3699) get_pg_deadlocks, get_pg_plan_capture_readiness and get_pg_index_bloat published truncated = rows.Count >= limit, the #3594 inference; the readiness and index tools withheld their summaries on the strength of it for pages that were the whole set, and get_pg_plans beside them Take(limit)'d a read already cut at limit and said nothing. McpHelpers.BoundPage turns (fetched at limit + 1, limit) into (page, truncated) in one place; the four tools are its first consumers and every count downstream is still a count of the page. McpPageContractTests missed the three because EveryPagedTool_ObservesTruncation walks the ten tools #3541 A3 named and its #3653 sibling walks one file - a roster gap, not a regex gap; the matcher catches both spellings and now witnesses the object-initializer one. NoTool_InfersTruncationFromItsCap_OnEitherSku sweeps every tool body on both SKUs with a population floor and fails on the three at dev. McpPayloadContractCensusTests pins the four unpinned rules as far as a census can honestly go: every catch returns through a shared shape (one ad-hoc site rostered) and the two shared shapes - a sentence and a JSON envelope - are an inventory for the vocabulary lane; every bounded parameter reaches a shared validator or one of five inline days_back refusals, and nothing clamps a parameter; every total_* assigned from a Count names a whole set (rostered), every latest-anchored reader projects its anchor (two known exceptions); truncation-key dialects, severity-word spellings and statement-terminal literal LIMITs are exact inventories. The DarlingMcpPgPlanToolsTests pin that handed the builder two rows at limit 2 and expected truncated encoded the inference as its own proof and becomes the boundary pair. Executed here through the built Darling.Tests.dll and against PostgreSQL 18.4 / TimescaleDB 2.28.1; positive controls: the three files at dev state fail the sweep by name, and BoundPage flipped to >= fails the live boundary, the builder pins and the helper theory.
This was referenced Sep 19, 2026
erikdarlingdata
added a commit
that referenced
this pull request
Sep 19, 2026
…le, Opus full tier (#3710) (#3718) * REVIEW-TRAPS.md: the repository's recurring defect shapes, one line each with the PR that bit (#3710) Seven layers (SQL, collectors and deltas, analysis, alerting and notifications, storage and MCP payloads, tests and censuses, workflows and CI, repository mechanics), fifty-odd lines, every one citing the pull request or issue that verified it: the >= limit truncation inference (#3594/#3679/#3699), ORDER BY 3 + 4 folding to a constant (#3613), MAX(max_dop) across plans (#3648), integer division in cntr_value_per_second (#3653 Q9), per-collection alert firing against a shorter cooldown (#3628/#3636), restart zeros in deltas and baselines (#3595/#3698/#3705/#3708), ScoreAll skipping amplifiers on a zero base and the non-virtual GetActiveEdges (#3688), the PgTargetFactKeys censuses (#3688), pg_total_relation_size vs the heap and the two TimescaleDB views that hide materializations and history (#3610/#3585), the -infinity crash backoff (#3629), the Slack block budget and the surrogate-pair cut (#3612/#3618/#3625/#3644), LCK_M_SCH_M folding into LCK so no SCH_M fact exists (#3709), one wait profile per run (#3709), and the CRLF round-trip. Each citation was read before it was written down; a wrong citation is worse than none. The review workflow reads this file before the PR body and must say, per trap the diff is near, whether it applies here and why. * Claude review: blind read before the body, callers rule, structured verdict with a ledger, and an Opus/Sonnet tier by changed path (#3710) The prompt runs in two phases. Phase 1 forbids gh pr view / git log / git show until the reviewer has read the diff, every changed region in context, the callers of every changed public symbol (Grep, both SKUs and both test projects) and .github/REVIEW-TRAPS.md, and written its own account. Phase 2 reads the body and grades every claim VERIFIED / UNVERIFIED / REFUTED, then reports the DIVERGENCE between the two accounts. The verdict is a fixed block -- Claims, Riskiest lines with the pinning test or UNPINNED, Traps near this diff, Divergence, Reviewer checklist -- ending in one machine-readable ledger line, posted through gh pr review --body-file - with a quoted heredoc (no Write tool; the allowlist is unchanged and still read-only). The #3650 verdict check now fetches the newest verdict in the window by id and fails the job when the body lacks the "## Verdict:" header or the ledger, or when the ledger's verdict disagrees with the review's API state (a "CHANGES REQUESTED" body submitted with --comment is the shape the guard would wave through). The guard's own arm is untouched and still keys on the state. A dorny/paths-filter@v4 step (the repo's pin style) routes changes under the shared libraries, Darling Storage/Analysis/Service, Lite Services/Mcp/Database, install/ and every .sql file to the full tier -- claude-opus-5[1m], 60 turns -- and everything else to claude-sonnet-5 at 30 turns; one decision step emits tier/model/max_turns, one review step consumes them (two steps would double the "review posted" accounting the guard does), and the prompt carries the tier and model into the ledger. A filter that dies degrades to FULL, not light. Step name "Claude review" and the ALWAYS POST sentinel are unchanged: both are read by claude-review-guard.yml. * Review tier: Alerting and Notifications read at the full tier (#3710)
16 tasks done
This was referenced Oct 11, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Six MCP pages now say what bounded them, so a limited page stops passing for a complete answer
Partial for #3541 — checklist item A3 "Hidden caps published as window totals". Does not close the issue.
Six tool groups on both SKUs carried a hidden reader cap:
LIMIT 200/50/500(or no bound at all) inside the SQL, under a tool that advertisedlimit, applied it withTake(limit)on top, and then published the capped row count under atotal_*name. A reasoning agent holding only the JSON read a 200-row page of a 5,000-event window as the window. The "widenhours_back" advice in the empty branches could never help, because the cap was on rows, not time. #3287 fixed exactly this shape onget_collection_logand established the pattern; this PR applies that one pattern, as one dialect, to every tool the finding named:get_blocking(Lite:get_blocked_process_reports)limit" was quietly hollow past the newest 200 rowsget_deadlocks/get_deadlock_detailget_alert_historydismissed = FALSEfilter — acknowledged criticals vanished, unstated anywhereget_long_query_completionsget_plan_correctionshours_backsaidget_waiting_tasks/get_wait_statslimitup to 1,000get_waiting_tasksshipped a bare envelope: server and rows, no window, no count, no boundThe one pattern
limit, bound as a SQL parameter (LIMIT $4), never a literal. Every touched reader const is pinned to end in a parameterisedLIMIT.limit + 1and setstruncatedonly when the extra row came back.count >= limitcannot tell a window holding exactlylimitrows from a busier one, and every test here asserts the pair —limit = N-1truncated,limit = Nnot — over N seeded rows.get_collection_logestablished:oldest_returned_<stamp>/newest_returned_<stamp>, plusordernaming the ordering. Under newest-first the oldest stamp is the reach of the read; underget_long_query_completions' duration ranking the stamps bound the slowest runs and say nothing about reach — the description says which case each tool is.total_*.total_events/total_deadlocks/total_alerts/total_completions/total_recommendations→*_returned.get_alert_historypublishesdismissed_excludedanddismissed_excluded_count(a realCOUNT(*)of what the default filter removed in the window), gainsinclude_dismissed(appended afteras_of, default false = the grid's own read), labels each rowdismissed, and an all-dismissed window now says "N dismissed alert(s) were excluded — re-run with include_dismissed" instead of calling itself quiet.get_blocking/get_deadlocksfetch up toFingerprintScanCeiling(2,000; lineage in the const's comment: a scan row carries the graph or both SQL texts, ten times the old cap, half of what the analysis pair-row readers fetch without XML) when adedup_keyis supplied, publishrows_examined/scan_truncated, and a no-match answer on a scan that ran out appends that fact and the remedy (anchoras_ofat the alert time). Without a key the scan is the page plus its sentinel row.get_long_query_completions: a completions tool sorted by duration keeps the window's SLOWEST. Lite gained a duration-ranked SQL read (GetSlowestLongQueryCompletionsAsync,ORDER BY duration_microseconds DESC NULLS LAST) for the tool; the grid keeps its chronological read. Both descriptions now say "the window'slimitSLOWEST, not its newest".truncatedappears in the tool description and in thelimitparameter's description, on both SKUs, and the Instructions tables carry the same sentence.get_deadlock_detailandget_blocked_process_xmlmoved their graph/XML predicate into the SQL (RecentDeadlocksWithGraphSql,BlockedProcessReportsWithXmlSqlon Darling;graphOnly/xmlOnlyon Lite's readers), solimitcounts graphs rather than rows the tool would have discarded, and the XML tool reads the XE arm alone (a DMV snapshot never carries a report).Why this shape, and what was rejected
COUNT(*)for the row-level tools: a true windowed count is a second scan per call on the store's largest edge tables for a fact the caller only needs as a boolean. The count is run for the alert-history filter, because "2 dismissed rows were hidden" is the disclosure that makes the filter honest and it is a cheap indexed count.get_active_queries' unboundedtotal_snapshotsbesideshownwas left alone — that is a real total, and its filter semantics are A13's item.cap + XE rows in hand, because the merge drops one DMV row per (pair, minute) an XE row already covers and the DMV collector samples once per cycle; fetched at the bare cap, a surplus made of XE-covered rows would vanish in the merge and a full page would read as complete. Both SKUs, same reason, in the reader remarks.include_dismissedrather than dropping the filter: dismissal is a Viewer operator's acknowledgement ("hide it from the grid"), which is the right default for a person at the grid and says nothing about whether the alert fired. The default stays the grid's; the payload now says so and measures it. On Lite, an alert dismissed after aging into the parquet archive is removed by the archive view itself and can be neither returned nor counted — the description says so.LocalDataServicereads gained a trailinglimit(defaulting to the grid's 50 / 200) so every WPF caller reads exactly what it always read; Darling's readers takecapand the Viewer has its own SQL.McpHelperspage helper. The one-dialect guarantee is held by census tests instead:PerformanceMonitor.Common/Mcp/McpHelpers.csis outside this lane's file boundary. Reported to the coordinator as the natural follow-up if the contract item wants a helper.Blast radius
total_events/total_deadlocks/total_alerts/total_completions/total_recommendationsare gone from these six payloads (renamed). Swept the web viewer (wwwroot/js) and Lite: every consumer keys on the array (events,deadlocks,alerts, …), none on the totals./api/read/*dispatch passesas_ofby name, so the appendedinclude_dismissedis invisible to it (it cannot yet be asked for; noted below).DarlingAlertReader.GetAlertHistoryAsynckeeps its signature (the triage endpoint calls it); the tool reads through the newGetAlertHistoryPageAsync.AlertHistoryReadRow/ LiteAlertHistoryRowgainDismissed.200 + XE rowsDMV candidates into the merge before re-capping to 200 — strictly closer to "the newest 200 merged" than before.get_wait_stats: only theLIMIT 50→$4binding pluswait_types_returned/truncated(the twin port of Lite's named defect). The A7 page-scoped-percent item is untouched.deprecated/Dashboard/Mcpcarries the sametotal_events/total_alertspage counts, but its caps live in its own SQL-Server-side readers outside this lane's boundary; not touched, reported.What it does NOT do
COUNT(*)to the row-level tools.include_dismissedthroughDarlingWebEndpoints'/api/read/get_alert_history(file out of boundary).Test plan
New, compile-verified here (the test projects are
net10.0-windows), first executed in CI:Lite.Tests/McpPageContractTests.cs— runs the real tools against a real DuckDB: the boundary pair for every tool (limit = N-1truncated /limit = Nnot), page bounds equal the seeded stamps,get_deadlock_detail/get_blocked_process_xmlcount graphs not rows, the merged blocking arms stay observable,get_alert_historystates/measures/lifts the dismissed filter and its all-dismissed empty branch names it,get_long_query_completionsreturns the window's slowest (seeded as the OLDEST) atlimit = 1and is duration-ordered,get_waiting_taskscarrieshours_back,get_wait_statsbinds the cap;AssertPagerefuses anytotal_*key; reflection pinstruncatedin every touched description andinclude_dismissedas an appended optional.Darling/Darling.Tests/McpPageContractTests.cs— the census: every paged Darling read bindsLIMIT $Nand no literal; every paged tool BODY on both SKUs has nototal_* = x.Count, no>= limitinference, and aCount > limitover alimit + 1fetch; the same tool name emits the same page-contract keys on both SKUs (the naming-drift pair compared on bounds/order); descriptions on both SKUs saytruncated/limit; the fingerprint readers namerows_examined/scan_truncated; discriminators witnessed against defect-shaped and fixed-shaped literals. PlusMcpPageContractLivePostgresTests(gated onDARLING_TEST_PG) running the same boundary pairs against live Postgres.DarlingMcpBlockingToolsTests(LIMIT $4, the two with-XML/graph consts share a body, scan ceiling bounds, live read assertsevents_returned),DarlingMcpAlertToolsTests((dismissed = FALSE OR $5), the count consts unbounded,include_dismissedcontract, live read asserts the new fields),DarlingMcpSessionToolsTests/DarlingMcpDataToolsTests/DarlingMcpPlanCorrectionToolsTests(LIMIT $4, literal gone),PlanCorrectionFrameLiveTests(newcaparg).Verified here: all four projects build with 0 warnings (
Darling.Service,LitewithEnableWindowsTargeting, both test projects). The Darling SQL was exercised end-to-end against a throwawaytimescale/timescaledb:2.28.1-pg18container (migrated withPgMigrations, rows planted, every touched tool called through its real method):LIMIT $4binds, the(dismissed = FALSE OR $5)bool parameter works on both scopes, the with-XML/graph consts page correctly, truncation flips exactly at the boundary, the dedup-key path reportsrows_examined, the slowest-first page holds the window's slowest (the oldest row), and the all-dismissed window hits the new empty branch. Container torn down. The Darling census regexes were emulated against the actual sources and pass.Which component(s) does this affect?
Checklist
dotnet build -c Debug)