Skip to content

get_ag_health fills its fleet-wide answer to the 32 KB budget by size, not a fixed group count - #4568

Merged
erikdarlingdata merged 4 commits into
devfrom
fix/ag-health-byte-budget
Sep 28, 2026
Merged

erikdarlingdata merged 4 commits into
devfrom
fix/ag-health-byte-budget

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Refs #4471 #4474

Why

A 42-group production fleet returned 63,333 characters against the shared 32 KB MCP response budget — about 2x over. The default page cap (11 groups) was sized from a fixture averaging about 2.7 KB per group (3 databases per AG); real groups averaged closer to 6 KB (2 replicas plus 6-14 database rows), so the fixed count overshot badly on real data.

What changes

DarlingAgReader.Build now fills the returned page by MEASURED serialized bytes, most-severe-first, instead of a fixed group count:

  • Groups are added one at a time in severity order; the walk stops before the group that would push the running UTF-8 size (envelope plus every group already added) past McpResponseBudget.DefaultBytes (32 KB) — never after.
  • At least 1 group is always returned, even a single group whose own bytes exceed the budget alone.
  • An explicit limit argument still applies as an upper bound underneath the budget (whichever is smaller cuts the page).
  • groups_returned, groups_total, and groups_truncated stay truthful against the whole scope. groups_truncated_note now names the actual reason — the byte budget or the caller's limit — since the two can now disagree.
  • Each candidate group is serialized exactly once (its own byte count is measured and reused), so the fill stays O(n), not O(n²), on a large fleet.
  • The fill is then checked against the response actually returned: the final result, note included, is serialized and measured, and while it is over the budget (and more than one group is left) the last group is dropped and the note rebuilt with the final counts.
  • Updated the DefaultGroupLimit doc comment and the tool's Description/parameter text to describe the new "as many as fit ~32 KB, most-severe-first" default, with limit still an upper bound.

Test plan

Pure pins added to Darling/Darling.Tests/DarlingAgReaderTests.cs (no live store needed), against a realistic wide-group fixture (2 replicas + N databases, real-shaped LSNs/queue sizes/suspend reasons):

  • 42 wide groups (10 databases each): fits ≤32 KB, groups_returned < 42, groups_truncated true, note names the byte budget.
  • 5 small groups (3 databases each): every group still comes back, no regression on the common small-fleet case.
  • 1 oversized single group (400 databases): exactly 1 group returned, never 0, and groups_truncated is false (a lone group isn't itself a cut).
  • Explicit limit smaller than what the byte budget would allow: caps at exactly that limit.
  • Groups that fit the budget without the truncation note but not with it: the final re-measure drops a group so the response, note included, fits (over the budget with that check disabled).
  • The live get_ag_health test with limit at or above the group count asserts the truthful counts, the budget wording in the note, and that the serialized response itself fits the budget.

Runtime RED on dev, measured: with the fill reverted to dev's count-based groups.Take(limit), the 42-wide-group pin returned 280,142 bytes (all 42 groups; the pin calls Build without the tool's default count) against its <= 32,768 assertion, so it fails on dev. Restoring the byte fill turned it green.

Also re-ran and confirmed green (no regression):

  • DarlingAgReaderTests (53 existing + 4 new = 64 total incl. surface/budget classes below)
  • DarlingMcpAgToolsSurfaceTests (tool surface/schema contract)
  • McpToolsListBudgetTests (tools/list byte ceiling — trimmed the tool's Description/parameter text to fit; the param get_ag_health.limit ceiling line dropped from 155 to 146, banking the saving as the test requires)

DarlingMcpAgToolsLivePostgresTests (gated on DARLING_TEST_PG, runs in CI) was not run here (no rig); it wasn't touched and its assertions (default-limit-caps behavior on a 42+1 group live fixture) still hold under the byte-fit fill since that fixture's groups are far under 32 KB — CI decides it.

Build: dotnet build Darling/Darling.Tests/Darling.Tests.csproj -c Release -p:EnableWindowsTargeting=true — 0 warnings, 0 errors.

CHANGELOG

SECTION: Fixed
ENTRY: - get_ag_health keeps its fleet-wide answer within the 32 KB MCP budget ([#4568]) - The fleet-wide default page now fills most-severe-first by measured size instead of a fixed 11 groups. On a real fleet whose groups carry many databases, those 11 groups came to about twice the budget. limit still caps the page, and groups_truncated says whether the size budget or the limit cut it.
REF: [#4568]: #4568

erikdarlingdata and others added 4 commits September 28, 2026 09:32
… not a fixed count

A 42-group production fleet (2 replicas plus 6-14 databases per group)
measured 63,333 characters against the shared 32 KB MCP response budget,
about 2x over. The fixed default page (11 groups) was sized from a
2.7 KB/group fixture; real groups ran closer to 6 KB.

The fill is now measured by actual serialized bytes: groups are added
most-severe-first while the running UTF-8 size stays under the budget,
stopping before the group that would cross it. At least one group is
always returned, even an oversized one. groups_truncated and its note
now say whether the byte budget or the caller's limit was the reason.

Refs #4471 #4474
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant