Repository navigation
MCP budget: default get_query_store_regressions under the response budget (#4198) - #4264
Merged
Merged
Conversation
Checkpoint: rig measured a 40-row/~4KB-text fixture at 25,527 bytes after the fix (was the documented 211 KB worst offender at old defaults). Darling and Lite both cut query_text to a 240-char preview behind full_text, round the ten numeric fields to 2 decimals, and drop the default row limit from 50 to 30. Web viewer row pinned to keep today's full-text behavior. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…s-default # Conflicts: # Darling/Darling.Tests/McpPayloadContractCensusTests.cs # Darling/Darling.Tests/McpToolsListBudgetTests.cs # Lite.Tests/McpToolsListBudgetTests.cs
…CutKeys Deltas: audit+deadlock+plan_corrections (prior dev merges) + qs_regressions TR (+214 Darling, +203 Lite). query_text_truncated merges plan_correction and qs_regression files into one entry. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
erikdarlingdata
marked this pull request as ready for review
September 25, 2026 08:58
erikdarlingdata
enabled auto-merge (squash)
September 25, 2026 08:58
Darling: HEAD=172,434 (qs_regressions +214) + dev=172,367 -> 172,581 Lite: HEAD=90,474 (qs_regressions +203) + dev=90,418 -> 90,621 Census: merge qs_regressions files into query_text_truncated + keep top_query_text_truncated Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
This was referenced Sep 25, 2026
erikdarlingdata
added a commit
that referenced
this pull request
Sep 25, 2026
Measured with McpToolsListBudgetTests after merging origin/dev (#4261, #4258, #4265, #4267, #4264, #4266 and #4268) plus this PR's catalog changes. Budget, census and tool-guide classes: 219/219. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
erikdarlingdata
added a commit
that referenced
this pull request
Sep 25, 2026
…nder the response budget (#4198) (#4272) * MCP budget: describe_custom_view_catalog default groups measures by source (#4198) Default default-argument call was 98,173 bytes (#4198's own measurement), three times the tool's 32 KB budget. It is pure static reference data (no server/store read), so the cut groups the 179 measures by source and keeps only key/displayName/kind/unitFamily/validAggregates per measure; source=<name> drills into one source's full detail, full_detail=true returns the original shape unfiltered. /api/catalog (the web Custom Views editor) calls the underlying builder directly, never this MCP method, so it is unaffected - pinned in DarlingComposeTests. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3 * Trigger CI (draft PR skipped the Build workflow) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ * Add comment to DarlingMcpCustomViewCatalogSizeTests.cs to trigger Build CI GitHub did not fire pull_request events for this draft-opened PR. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ * fix(#4272): reword validAggregates description - compact drops allowedDimensions, source=<name> returns it Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3 * #4272: set TotalCeilingBytes to the measured 174,236 after merging dev Measured with McpToolsListBudgetTests after merging origin/dev (#4261, #4258, #4265, #4267, #4264, #4266 and #4268) plus this PR's catalog changes. Budget, census and tool-guide classes: 219/219. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
erikdarlingdata
added a commit
that referenced
this pull request
Sep 25, 2026
…as merged After merging dev, the per-tool fixes for get_blocking (#4267), get_collection_log (#4265), get_query_store_regressions (#4264), get_collection_health (#4268), describe_custom_view_catalog (#4272) and get_query_store_top (#4273) are all in, and get_fleet_overview already fit (CI's stale-exemption message). Every row goes; CI's live run decides. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
3 tasks done
erikdarlingdata
added a commit
that referenced
this pull request
Sep 25, 2026
* Add the #4198 MCP read-tool budget pin (part of #4198) McpReadToolBudgetLiveTests reflects over every [McpServerTool] method on every [McpServerToolType] class in the Darling service assembly, excludes write tools/analyze_*/compare_*/audit_config, binds each tool's DI services and server_name generically, and asserts the reply stays under McpResponseBudget.DefaultBytes unless the tool is named in an explicit, byte-stamped exemption roster. A fix lands by deleting its row. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3 * MCP budget pin: empty ExemptOffenders now every #4198 per-tool lane has merged After merging dev, the per-tool fixes for get_blocking (#4267), get_collection_log (#4265), get_query_store_regressions (#4264), get_collection_health (#4268), describe_custom_view_catalog (#4272) and get_query_store_top (#4273) are all in, and get_fleet_overview already fit (CI's stale-exemption message). Every row goes; CI's live run decides. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ * MCP budget pin: get_collection_health is still over on this fixture CI on 08e6e81 measured get_collection_health at 35,145 B, over the 32 KB default. #4268 compacts only healthy collectors with nothing to report, and this fixture's collectors are not healthy, so nothing compacts. The row goes back, under #4198, which stays open for it. Every other exempt tool fits. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #4198.
Why
get_query_store_regressionshad no response-size budget.McpResponseBudget.csflagged it as the worst #4198 offender: 211 KB at default arguments on a busy production store, over 6x the 32 KBMcpResponseBudget.DefaultBytestarget. Two things drove that size.query_textis an unbounded Query Store text sample, one per row, at up to 50 rows by default. The ten numeric fields (durations, CPU, reads, regression percents) come straight offAVG()over microsecond integers. They serialize with a long, meaningless decimal tail instead of a display-rounded one, and that weight repeats across ten fields on every row.What changes
Darling (
DarlingMcpQueryStoreRegressionTools.cs) and Lite (McpQueryTools.cs) both change the same way:query_textis a 240-character preview by default (the same lengthget_store_query_statsalready usesfor its own
full_text), with a newquery_text_truncatedflag. A newfull_textargument (defaultfalse) opts back into the whole text, the MCP read tools have no default response-size budget: at default arguments 12 tools return >50 KB and 3 return >100 KB for one server, more than an agent client's per-result cap #4198 precedent fromget_store_query_stats.limitdrops from 50 to 30. This is sized from the measured bytes per row so a default callstays well under budget instead of sitting right at its edge.
desktop viewer's own read are untouched. Every existing pinned test uses whole-number fixtures that round
trip unchanged (verified below).
/api/readrow forget_query_store_regressionsnow passesfull_text: QueryBool(c, "full_text", true)in the dispatch table, and the catalog row getsPBool("full_text", true). That keepsthe viewer showing today's full query text (it renders it via
codeDisclosureon the Query StoreRegressions tab). No
wwwroot/jschange was needed:server-tabs.js's call never sentfull_textexplicitly, so the row's own default carries it. A new no-rig test,
QueryStoreRegressionsWebDefaultTests.cs,pins this.
FieldPreviewCutKeys(census) gets aquery_text_truncatedentry, added alongside the deadlock-detaillane's own entry. Both merged cleanly from
origin/dev.McpToolsListBudgetpins are updated on both products.TotalCeilingBytesis raised by exactly themeasured growth (+214 Darling, +203 Lite), stacked on top of the concurrent deadlock-detail lane's own raise
after merging
origin/dev.This PR does not touch
DarlingQueryStoreRegressionReader's SQL, or Lite's equivalent reader, per the #4198note for this tool. It only changes how the tool projects the read into JSON.
Test plan
Rig: reused the stopped
rig-taPostgres rig (port 55976, already UTC), started in the background, confirmed"ready to accept connections", then dropped and recreated
darlingtestandprobe.DarlingMcpQueryStoreRegressionsBudgetLiveTests(Darling): plants 40 regressed querieswith ~4,091-character query text each. That fixture is the same order of magnitude as the documented
211 KB at 50 rows before this fix. After the fix: 25,527 bytes for the default call (30 rows
returned,
truncated: true), under the 32,768-byte budget.full_text: truereturns the whole text andno
query_text_truncated.QueryStoreRegressionsBudgetTests(Lite.Tests, no rig): same fixture shape, sameassertions. It passes.
DarlingQueryStoreRegressionsSurfaceAndSqlTests.ParamContract_AllOptional_MatchesLiteupdated for thenew
full_textparameter.McpPayloadContractCensusTests: myFieldPreviewCutKeysentry plus the deadlock-detail lane's, bothpresent after the merge.
McpToolsListBudgetTests(both products): pins andTotalCeilingBytesupdated to the measured totals.origin/dev: Darling.Tests, the classes this PR touches plusMcpZeroIsAMeasurementTests. 106 passed, 0 failed. Lite.Tests, the classes this PR touches.9 passed, 0 failed.
Darling.Testssuite: not run. Every class this PR's diff touches ran targeted and green (listedabove), but the full suite must run before arming auto-merge.
Lite.Testssuite: not run, same reason.Rig stopped (
pg_ctl stop -m fast) after the last test run.CHANGELOG entry
SECTION: Fixed
ENTRY:
get_query_store_regressionsdefault calls stayed under the MCP response budget ([MCP read tools have no default response-size budget: at default arguments 12 tools return >50 KB and 3 return >100 KB for one server, more than an agent client's per-result cap #4198]). Theworst-offending read measured 211 KB at default arguments on a busy store. It now previews query text to 240
characters with a
full_textopt-in, rounds its numeric fields, and defaults to 30 rows instead of 50. Theweb viewer is unaffected: it keeps the full query text it always showed.
REF:
[MCP read tools have no default response-size budget: at default arguments 12 tools return >50 KB and 3 return >100 KB for one server, more than an agent client's per-result cap #4198]: MCP read tools have no default response-size budget: at default arguments 12 tools return >50 KB and 3 return >100 KB for one server, more than an agent client's per-result cap #4198
Generated with Claude Code
https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3