Repository navigation
Add get_store_host MCP read for the store host profile (#4214) - #4282
Conversation
get_store_host reuses part 1's DarlingStoreHostProfile.GatherAsync (#4271) to serve the host/store/settings profile over MCP and /api/read, with its own enclosing deadline (store-size reads scale with the store, #3199) and a fixed #4271 CLI test that assumed a bootstrapped rig's own postgresql.conf and "darling" database rather than building its own. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
get_store_host is the 160th Darling MCP tool; DarlingMcpInstructions.cs already carried this bump but README.md and llms.txt did not, failing CrossAppMcpToolInventoryPinTests.RootReadmeDarlingToolCensus_MatchesTheScannedInventory and LlmsTxtToolCensus_SpansTheTwoCurrentEditions in CI. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…vilege roles DarlingMcpStoreHostToolsLiveTests calls DarlingMcpStoreHostTools.GetStoreHost through the SAME least-privilege roles it runs under in production (mcp for the MCP host, viewer for the web /api/read row), not the rig superuser every other #4214 part 2 test uses. Proves all four verdicts (matches, stale_after_hardware_change, operator_override, not_managed) are reachable under each role: - work_mem: a per-role ALTER ROLE ... SET default applied before the role's first connection, matched against this host's own DeriveMemorySettings figure -> matches (shared_buffers/max_connections are PGC_POSTMASTER and the live rig would need a restart to move them, so work_mem, PGC_USERSET, is the one setting a live value can be set for without touching the rig's real conf or going through postgresql.auto.conf, which would misattribute to operator_override). - max_connections: a managed-block assignment of 100 against the TargetMaxConnections=200 constant -> stale_after_hardware_change. - maintenance_work_mem: left unassigned in the fake managed conf -> operator_override (ClassifyVerdict's default arm). - A second gather with Managed=false -> not_managed for every setting. No permission gap found: pg_extension/pg_stat_database/pg_settings carry no Darling-authored grant anywhere in DarlingManagedRoles, and this proves PostgreSQL's own PUBLIC defaults are what let mcp/viewer read them; the one grant that matters (SELECT on collect, for timescaledb_information.chunks visibility) already exists in production provisioning and is mirrored here. Ran twice against the rig-th rig (port 55983): 1 total, 0 failed both times; confirmed host_mcp_test/host_viewer_test roles are dropped after each run. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
McpToolsListBudget/DarlingMcpStoreHostTools.txt pins the new tool's served block (510 chars) and McpToolsListBudgetTests.TotalCeilingBytes moves 173,157 -> 173,773 (+616, the tool's own JSON scaffolding included) to match: get_store_host has no parameters, so this is the tool's whole entry. DarlingMcpStoreHostBudgetLiveTests proves the response-size half: get_store_host's payload is a fixed shape (host/store facts plus one row per sizing-relevant setting, 8 today), not row-scaling like the other #4198 tools, so one live call in the not_managed shape (the longer of the two source strings) is the whole proof. Passed at well under McpResponseBudget.DefaultBytes (32 KB) against the rig-th rig. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
Adds a "Store host" section to the Fleet Sweeps page, below Watch items: reads get_store_host through /api/read (readTool, the page's existing read pattern - no raw fetch, no second copy of the verdict math #4214 part 1 already owns) and renders host facts (platform, CPUs, RAM, data volume, managed flag, PostgreSQL/TimescaleDB facts) plus a per-setting table (current/derived/source/verdict), with every stale-* verdict highlighted using the page's existing sev-Warning/ sev-Healthy vocabulary. Read-only, like the rest of this page (no apiSend call) - the CLI --check-settings verb and the companion sizing issue own any write. Uses the page's own local table()/cellText() helpers (no new table implementation) and four new util.js imports (fmtMb, fmtPct, fmtBool, fmtText) already used elsewhere in the app. FleetSweepWebFeedTests gains a source pin (no JS runner in this pipeline) proving the section title, the get_store_host read call, and the startsWith("stale") highlight predicate. Verified sev-Warning/ sev-Healthy are real table.data td rules in app.css, and node -c confirms the edited file parses. FleetSweepWebFeedTests (59) and DarlingWebAssetsTests all pass. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…t-read # Conflicts: # Darling/Darling.Tests/McpToolsListBudgetTests.cs
Full suite (13,992 tests) surfaced two defects caused by this branch:
- DarlingMcpStoreHostBudgetLiveTests reached the shared DARLING_TEST_PG
store without [Collection("live-postgres")], failing
LivePostgresCollectionHygieneTests. Added the attribute, matching
its sibling DarlingMcpStoreHostToolsLiveTests.
- DarlingMcpFleetSweepToolsTests.TheMirrorsLoggerSeat_... pinned the
old BuildReadDispatch(logger) call-site text; BuildReadDispatch now
also threads postgresConfig (get_store_host's config seat) per this
PR's earlier commit, so the literal moved to
BuildReadDispatch(logger, postgresConfig) - the same kind of pin
update already made for FleetSweepWebFeedTests/SharedBaselineCacheTests,
missed for this third site.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…t-read McpToolsListBudgetTests: keep every change-log line. TotalCeilingBytes re-measured on the merged tree: dev after #4272 and #4273 (174,373) plus get_store_host's 616 bytes = 174,989. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
Round-1 security review at f71f154Scope: the diff at f71f154 and the part-1 code that it calls ( Result: no High findings. There are two Medium findings, four Low findings, and one pre-existing Low finding that this PR does not widen. We recommend fixing both Medium findings in this PR, because each fix is small. Answers to the six questions
FindingsMedium 1: absolute conf-file paths reach every remote seat
Medium 2: an uncached, unthrottled store-scaling probe that the UI runs again every minute
Low 3: the whole PostgresConfig, owner secret included, is now injectable
Low 4: no test pins that PostgresConfig stays out of the served schema
Low 5: the least-privilege proof does not cover the store-fact reads
Low 6: hardcoded password for the test login roles
Low 7 (pre-existing, not widened here): error bodies carry
|
Round-1 fix lane: no code landed, full plan below for the next laneThis session hit the context watchdog's hard block (past ~300k transcript tokens) during All of the time went into reading the affected files and getting a stronger reviewer's Plan for the next lane, in order (validated by an advisor pass over the full orientation)Item 1 (Medium 1, conf paths).
Item 2 (Medium 2, cost) — the settled cache design. Do NOT use a bare static field inside the
Item 3 (Low 3, trimmed PostgresConfig). Two call sites, both currently pass
Item 4 (Low 4) — two independent halves.
Items 5/6 (Darling.Tests/DarlingMcpStoreHostToolsLiveTests.cs).
Item 7 (description tail + budget). All three edits land strictly after
sweeps.js (small-fix version exists — do this, not the "leave the poll" fallback).
Process note for the coordinatorThe brief's What's left (everything)No code written, no tests added, no build run, no rig started. The plan above is validated by a |
Round-1 review Low 3: the full PostgresConfig (owner connection string included) was injectable via DI at both the MCP and web host wiring, even though get_store_host's GatherAsync only ever reads Managed/DataDirectory off it. Register a trimmed copy at both call sites instead, so a future [McpServerTool] that takes a PostgresConfig parameter cannot receive the owner secret through this seat. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
Round-1 review: three fixes to the description tail (strictly after <<GUIDE>>, so the served tools/list head is unchanged - confirmed by McpToolsListBudgetTests, 5/5 green with no fixture edit needed). - data_volume_note -> data_volume.note (the payload key is nested under data_volume, not a flat data_volume_note). - Add the missing "ready" field to the data_volume parenthetical list. - Document gathered_at (UTC) and the 5-minute shared cache ahead of item 2 landing the field/cache itself, per the brief's item ordering. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…settings.source Round-1 review Medium 1: get_store_host's settings[].source printed a full local filesystem path for a managed block or operator override, fine on the CLI's own shell but not fine handed to a remote MCP/web caller. - HostSettingProfile gains SourceFile/SourceLine (default null/0), set only by the managed-attribution branch of GatherSettingProfilesAsync. SourceDescription itself is untouched, so --check-settings and the startup log line still print the full path as before. - New DarlingStoreHostProfile.FormatSourceForMcp: passes SourceDescription through verbatim when there is no SourceFile (not-managed/unreadable - already a kind word, never a path), otherwise renders "<kind> (<path>: <line>)" with the path redacted to a data-directory-relative path when inside the managed data directory, or the bare file name only when outside it (never a leading ../ that would still leak the parent tree). - get_store_host resolves the managed data directory once and reuses it for every settings row via the new formatter, in place of the raw SourceDescription. Tests: 3 new pure-function tests on FormatSourceForMcp (inside data directory, outside it, no SourceFile) plus the existing DarlingStoreHost- ProfileTests/StartupHostProfileLogTests/McpPayloadContractCensusTests/ McpToolsListBudgetTests classes, 124/124 passing. Revert-proved: gutting FormatSourceForMcp to return SourceDescription unconditionally failed the two new sanitization tests; restored and re-verified green. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…d data directory Round-1 review Low 4, two independent halves. Census: a new McpServiceParameterDiSeatCensusTests pins that every DI-service-typed [McpServerTool] parameter (NpgsqlDataSource, PostgresConfig, DarlingAnalysisService, ILogger today) has an AddSingleton<T>/AddSingleton(new T seat in DarlingMcpHostService.cs's own source text. Without one, the SDK would serve that parameter as a client argument instead of resolving it from DI - for PostgresConfig, a remote caller could then set managed:true with a UNC dataDirectory. Coupled to item 3 by design: before item 3 trimmed the registration, AddSingleton(config.Postgres) did not name PostgresConfig in source text, so this test only started passing once item 3 landed (confirmed by reverting item 3's line, below). UNC refusal: new DarlingStoreHostProfile.TryResolveProfileDataDirectory wraps DarlingManagedPostgres.ResolveDataDirectory and refuses (returns null) a UNC-resolved path, mirroring DarlingManagedPostgres.TryResolveConfPath's own UNC refusal for an include directive - without it, a UNC-configured managed store would make the service open a conf file, or stat a volume, on a remote share as its own account (an NTLM-relay vector). Does not change ResolveDataDirectory itself, shared by the writer/provisioning path. Wired at both call sites: GatherSettingProfilesAsync's dataDirectory local (falls through to the same not-managed path a bring-your-own store already takes - never opens a file), and ResolveVolumeAnchor's managed arm (falls back to AppContext.BaseDirectory, same as its existing BYO branch). Tests: McpServiceParameterDiSeatCensusTests (new) plus 3 new pure-function tests on TryResolveProfileDataDirectory/ResolveVolumeAnchor. 135/135 passing (4 skipped, live-gated, no rig this pass). Revert-proved twice: reverting item 3's AddSingleton line failed the new census test; separately, gutting TryResolveProfileDataDirectory's UNC check failed both new UNC tests. Both restored and re-verified green. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…oles Round-1 review Low 6: DarlingMcpStoreHostToolsLiveTests.cs's two LOGIN test roles (each granted SELECT on every table in collect) used a hardcoded password that is public in this repository. Cleanup runs in finally, but a cleanup failure after a failed body is swallowed, and a killed run skips cleanup outright, so a role that survives on a reachable rig would be a known, reusable credential. Generate the password fresh per run instead (Convert.ToHexString(RandomNumberGenerator.GetBytes(16)) - hex output needs no escaping in the DDL string). Scope is this file only; DarlingSecuritySplitLiveTests uses the same pattern but is out of scope here. Build only (this is a live-gated test class, no rig this pass): 0 warnings, 0 errors; the class's own live Fact still skips cleanly without DARLING_TEST_PG set. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
Round-1 fix lane, second pass: 5 of 8 items landed, context watchdog stopped the 6th mid-designBranch Item 3 (Low 3, trimmed PostgresConfig) —
|
…UNC refusal GetStoreHost's mcpDataDirectory local called the raw DarlingManagedPostgres.ResolveDataDirectory, while GatherSettingProfilesAsync (Low 4) now refuses a UNC-resolved data directory via TryResolveProfileDataDirectory. The doc comment on mcpDataDirectory claimed it was "the SAME data directory GatherSettingProfilesAsync resolved" - no longer true under a UNC-configured managed store. Harmless today (every setting's SourceFile is null in that case, so FormatSourceForMcp never reaches the sanitizer), but switches mcpDataDirectory to the same TryResolveProfileDataDirectory call so the comment is accurate again and there is no divergent path. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
New StoreHostProfileCache (Darling/PerformanceMonitor.Darling.Service/StoreHostProfileCache.cs): a single immutable CacheEntry behind one volatile reference field (not a bare tuple field, which can tear under concurrent reads), a SemaphoreSlim gate for single-flight gather, and an injected clock. A cache hit never calls the gather delegate at all, so it never opens a connection. A failed gather is never cached: the finally always releases the gate, and _entry is only assigned after gather returns successfully. Wiring: GetStoreHost takes a third StoreHostProfileCache parameter; postgres.OpenConnectionAsync moved inside the gather delegate; the response gains gathered_at (naive-UTC, matching every other "_at" timestamp on an MCP payload in this codebase per DarlingFleetReader/DarlingAgReader/ DarlingMcpConfigHistoryTools/DarlingMcpTrendTools - not the DateTimeKind.Utc the original plan named). DarlingMcpHostService.cs registers the process-wide StoreHostProfileCache.Shared singleton via the typed-generic AddSingleton<T> overload McpServiceParameterDiSeatCensusTests requires. DarlingWebEndpoints.cs's direct-call dispatch passes the same Shared instance. StoreHostProfileCache is public, not internal as the handoff's draft had it: GetStoreHost is a public [McpServerTool] method, and a parameter type may never be less accessible than the method (CS0051) - matches the existing PostgresConfig/DarlingAnalysisService precedent. GetOrGatherAsync itself stays internal, since its signature carries the internal HostProfile type. Test call sites (DarlingMcpStoreHostBudgetLiveTests.cs, DarlingMcpStoreHostToolsLiveTests.cs) each get a fresh StoreHostProfileCache instance, never the production Shared singleton. The two calls in DarlingMcpStoreHostToolsLiveTests.cs's role loop (managed config then BYO config against the same data source) each need their OWN fresh cache: the cache holds one un-keyed entry, so sharing one across those two calls would serve the first call's cached profile back for the second regardless of which PostgresConfig was passed. New Darling/Darling.Tests/StoreHostProfileCacheTests.cs, 4 tests: two concurrent callers gather once (proven deterministically via a two-TCS choreography, no Task.Delay), a call inside 5 minutes gathers nothing, a call after 5 minutes gathers again, a throwing gather is not cached. Revert-proved: gutted TryGetFresh to always return null, confirmed both TwoConcurrentCalls_GatherOnce and ACallInsideFiveMinutes_GathersNothing fail, restored, rebuilt 0/0, reverified all green. McpServiceParameterDiSeatCensusTests' pinned service-parameter-type list grows from four names to five (StoreHostProfileCache added), with the reason in the test's own comment. Build: 0 Warning(s), 0 Error(s). Targeted run (StoreHostProfileCacheTests, McpServiceParameterDiSeatCensusTests, McpPayloadContractCensusTests, McpToolsListBudgetTests, DarlingStoreHostProfileTests, StartupHostProfileLogTests, DarlingCliCommandsHostCheckTests, DarlingMcpStoreHostToolsLiveTests, DarlingMcpStoreHostBudgetLiveTests, FleetSweepWebFeedTests, DarlingWebAssetsTests, SerialLoopStoreSizeSourceTests): 220 total, 0 failed, 6 skipped (live-gated, no rig this pass). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…IDisposable (CA1001) A non-incremental rebuild surfaced CA1001: the class owns a disposable SemaphoreSlim field (_gate) but was not itself IDisposable. Mirrors DarlingWebOidcClient's own gate disposal. Safe for the Shared singleton despite two hosts (DarlingMcpHostService, DarlingWebEndpoints.cs's direct-call dispatch) both holding a reference to it: both register it with the DI INSTANCE overload (AddSingleton<StoreHostProfileCache>(Shared)), and the built-in container never disposes an instance it did not itself construct, so no container shutdown disposes Shared out from under the other host. Build: 0 Warning(s), 0 Error(s) on a --no-incremental rebuild. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…ng on a poll tick
app.js's route() now forwards opts (the 60s poll's { poll: true } marker) to renderSweeps, which
forwards it to renderStoreHost. renderStoreHost is split into itself (fetch/cache) and a new
renderStoreHostPayload(box, p) (pure render), with a module-level lastStoreHostPayload. A poll
tick with an already-fetched payload replays it through renderStoreHostPayload instead of calling
readTool("get_store_host", {}) again; every non-poll call (first paint, hashchange, the span
control) always fetches fresh. The profile this card reports (host RAM/CPUs/disk, PostgreSQL/
TimescaleDB versions, sizing verdicts) changes on a hardware or version change, never per-tick, so
a 60-second poll paying a fresh network round trip for it was pure waste - the tool's own 5-minute
server-side cache (item 2) already made repeat calls cheap for the STORE, but not for the browser's
own network cost.
FleetSweepWebFeedTests.TheSweepPage_IsWiredIntoTheShellAndTheRouter pinned the literal
"renderSweeps(main)" call text; updated to "renderSweeps(main, opts)" to match.
Build: 0 Warning(s), 0 Error(s) (--no-incremental). Targeted run (StoreHostProfileCacheTests,
McpServiceParameterDiSeatCensusTests, McpPayloadContractCensusTests, McpToolsListBudgetTests,
FleetSweepWebFeedTests, DarlingWebAssetsTests, SerialLoopStoreSizeSourceTests,
DarlingCliCommandsHostCheckTests): 164 total, 0 failed, 4 skipped (live-gated, no rig this pass).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…e live test GetStoreHost_AsMcpAndViewerRoles_SurfacesAllFourVerdicts_Live gains three assertions per role, against the managed-config response: store.size_bytes and store.timescale_version are not null (both come off reads PUBLIC can run with no Darling-authored GRANT, so a regression there is a silent "unavailable" null rather than a thrown exception), and store.uncompressed_chunk_count matches a fresh owner-side read of the exact same query (DarlingStoreHostProfile. UncompressedChunkSizeSql, reused rather than duplicated) taken immediately after the tool call. Written per the handoff; not run live in this pass (no rig). Build: 0 Warning(s), 0 Error(s). The class still skips cleanly without DARLING_TEST_PG. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
Round-1 fix lane, third pass: all 4 remaining items landedBranch Item 1's open thread —
|
# Conflicts: # Darling/Darling.Tests/McpToolsListBudgetTests.cs
Rig pass: live classes and the full suite, onceBranch Worktree note
1. Merge origin/dev
175,170 = the shared #4273 base both sides branched from, plus this PR's own +616 ( No other file had real conflict markers ( Build: Merge commit 2. RigBuilt fresh at The worker numbers mirror CI's own 3. Live/gated classes (
|
| Class | Total | Failed | Skipped | Live-ran |
|---|---|---|---|---|
DarlingMcpStoreHostToolsLiveTests |
1 | 0 | 0 | 1 ran live |
DarlingMcpStoreHostBudgetLiveTests |
1 | 0 | 0 | 1 ran live |
DarlingCliCommandsHostCheckTests |
7 | 0 | 0 | 7 ran live |
StoreHostProfileCacheTests |
4 | 0 | 0 | n/a (deterministic unit tests, not rig-gated) |
McpToolsListBudgetTests |
5 | 0 | 0 | n/a (not rig-gated; run above during merge resolution) |
McpServiceParameterDiSeatCensusTests |
1 | 0 | 0 | n/a (not rig-gated) |
FleetSweepWebFeedTests |
39 | 0 | 0 | n/a (not rig-gated) |
DarlingWebAssetsTests |
21 | 0 | 0 | n/a (not rig-gated) |
Every rig-gated class ran its live test(s) with zero skips (confirms DARLING_TEST_PG /
DARLING_TEST_PGRUNTIME reached them) and every class passed clean.
4. StoreHostProfileCache.Shared watch
Checked both live call sites against the process-wide Shared cache risk the brief named:
Darling/Darling.Tests/DarlingMcpStoreHostToolsLiveTests.cs:119(managed-config role call) and:150
(BYO-config role call) each construct their ownnew StoreHostProfileCache(TimeSpan.FromMinutes(5))—
two separate instances, one per call, not.Shared.Darling/Darling.Tests/DarlingMcpStoreHostBudgetLiveTests.cs:56likewise builds its own fresh instance.
No cross-test or cross-role profile bleed observed; this matches what the prior lane's report already
described fixing. Nothing to change here — reporting the check, not a finding.
5. Fixes
None. Every class above passed on the first run; no code changes were needed in this PR's files.
6. Full suite, once
Darling.Tests.exe with no class filter, same two rig env vars, against the freshly recreated
darlingtest:
Total: 14076, Errors: 0, Failed: 0, Skipped: 29, Not Run: 1, Time: 948.584s (~15.8 min)
- 0 failures, 0 errors.
- 29 skipped: all env-gated live tests this rig doesn't set up (CSVLOG/JSONLOG/AUTOEXPLAIN clusters,
DARLING_TEST_PGRUNTIME_OLDfor the PG17-upgrade fixture) — expected, this rig only stands up the one
primary cluster the brief named. - 1 not run:
Darling/Darling.Tests/McpSchemaCompatServiceLeakRaceTests.cs:160,
[Fact(Explicit = true)], correctly excluded by the runner's default (-explicit off). Pre-existing,
unrelated to this PR's files. - Neither of CI flake: PgTarget anomaly/blocking worst-tile assertions fail on first attempts unrelated to the change #4274's known first-attempt flakes nor
TrendPayloadBudgetLiveTests' local stream error
appeared in this run, so there was nothing to isolate and re-run alone.
Installer.Tests was never run.
7. Rig teardown
pg_ctl.exe -D C:\GitHub\worktrees\rig-h2fix\data stop -> server stopped. Data directory left in place;
extracted runtime at C:\GitHub\worktrees\rig-h2fix left for reuse. No process killed by image name; only
the pg_ctl/postgres process this run started.
What the coordinator should double-check
- The merge commit (
9e3563d1) lacks the Co-Authored-By/Claude-Session trailer (see note above under
"Merge origin/dev"); it's the only commit this pass made. - The brief's placeholder session URL (
session_01TszxYhJJbTEh4LrZ56NYo3) is identical to the previous
lane's own session URL from PR comment 5835678714 — looks like a stale copy in the dispatch template. Not
consumed here since no fix commits were needed, but worth checking for other lanes on this same template. TotalCeilingBytes = 175_170is a growth-only ceiling; if a future PR on this branch changes any tool
head byte count, re-measure rather than trust the arithmetic above.
🤖 Generated with Claude Code
Part of #4214.
Why
Part 1 (#4271, merged) built the store host profile model, the
--check-settingsCLI verb and the--validate-configregistry fix. Nothing could reach that profile except a shell on the store's own host: noMCP read, no web panel. This PR adds the read half (
get_store_host) and the web Store host panel. Cloudidentity through IMDS (#4214's third piece) is a later, separate lane.
What changes
get_store_host(Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpStoreHostTools.cs): callsDarlingStoreHostProfile.GatherAsyncdirectly (part 1's own orchestration - no second copy, ruling 1).Returns platform/RAM/data-volume host facts, PostgreSQL/TimescaleDB store facts, and a verdict
(
matches/stale_after_hardware_change/operator_override/not_managed) per sizing-relevant setting, plusan
any_stalesummary flag. No parameters - a snapshot, store-level likeget_store_metrics, not a block onit (ruling 1). Reads as the least-privilege
mcprole (ruling 2); the managed conf-file attribution is alocal disk read by the service process itself, same as part 1.
ServiceCommandDeadlines.McpStoreHostProfileSeconds(25s), an outer linked CTSaround
GatherAsync. UpdatedSerialLoopStoreSizeSourceTests' Store host visibility: a cross-platform host profile, a --check-settings verb, and a get_store_host read, so "is this store sized right" has an answer without shell access #4214 allow-list entry for the second caller.PostgresConfigregistered as an MCP-host DI singleton;DarlingMcpStoreHostToolsregistered via
.WithGeminiCompatibleTools<T>().get_store_hosttoKnownLiteMissingMcpTools(Lite has no managed PostgreSQL store - a SKU boundary)./api/read: theR(CatOverview, ...)catalog row and the dispatch row inDarlingWebEndpoints.cs.BuildReadDispatch/ConfigurePipeline/MapAlleach gained an optionalPostgresConfig? postgresConfig = nullparameter (mirroring the existingloggerclosure pattern forget_sweep_reports). Updated the source-pin tests that literal-matched the old call text(
FleetSweepWebFeedTests,SharedBaselineCacheTests, and a third one this lane found -DarlingMcpFleetSweepToolsTests.TheMirrorsLoggerSeat_..., which still pinnedBuildReadDispatch(logger)after the signature grew a second argument).
was updated but these two files were not, failing
CrossAppMcpToolInventoryPinTestsin CI.McpToolsListBudget/DarlingMcpStoreHostTools.txt: the new tool's served-block pin (510 chars, noparameters).
McpToolsListBudgetTests.TotalCeilingBytesraised by the tool's own +616 bytes (merged withorigin/dev's own accumulated raises from other lanes; final value is the sum, not two deltas added by hand).
DarlingMcpStoreHostBudgetLiveTests: proves the payload stays underMcpResponseBudget.DefaultBytes-this tool's shape is fixed (host/store facts plus one row per setting, 8 today), not row-scaling like the
other MCP read tools have no default response-size budget: at default arguments 12 tools return >50 KB and 3 return >100 KB for one server, more than an agent client's per-result cap #4198 tools, so one live call in the
not_managedshape (the longer of the two source strings) is thewhole proof.
DarlingMcpStoreHostToolsLiveTests: a live test that callsget_store_hostthrough the SAMEleast-privilege roles it runs under in production -
mcp(the MCP host's own connection role) andviewer(the web
/api/readrow's connection role,TryBuildViewerConnectionStringFromStoredCredential) - ratherthan the rig superuser every other Store host visibility: a cross-platform host profile, a --check-settings verb, and a get_store_host read, so "is this store sized right" has an answer without shell access #4214 part 2 test uses. Proves all four verdicts are reachable under each
role. No permission gap found:
pg_extension/pg_stat_database/pg_settingscarry no Darling-authoredgrant anywhere in
DarlingManagedRoles- PostgreSQL's own PUBLIC defaults are what letmcp/viewerreadthem, and this test is what actually proves that rather than trusting it. The one grant that does matter
(SELECT on
collect, fortimescaledb_information.chunksvisibility) already exists in productionprovisioning and is mirrored in the test's own disposable roles.
host" section on the Fleet Sweeps page (
wwwroot/js/pages/sweeps.js), below Watch items - not a new Fleetpage card, which is what the earlier draft of this PR proposed before finding no existing "store metrics
panel" page to extend. Reads
get_store_hostthrough/api/read(readTool, the page's own readpattern), renders host facts plus a per-setting table (current/derived/source/verdict), with every
stale-*verdict highlighted via the page's existingsev-Warning/sev-Healthyvocabulary. Read-only, likethe rest of the page (no
apiSendcall - the CLI verb and the companion sizing issue own writes). Uses thepage's own local
table()/cellText()helpers, no new table implementation.FleetSweepWebFeedTestsgained a source pin (no JS runner in this pipeline) for the section title, the read call, and the
startsWith("stale")highlight predicate.DarlingCliCommandsHostCheckTests. CheckSettingsAsync_ManagedStoreWithAStaleBlock_ReturnsStaleSettingsExitCode_Gatednow creates and drops itsown
darlingdatabase and uses a throwaway temp conf directory instead of the rig's real one. Confirmedlive: it was failing with
StoreUnreachableon a hand-built rig that never has adarlingdatabase.Confirmed root cause (ruling 7)
DarlingManagedPostgres.DatabaseName = "darling"andUserName = "darling"are hardcoded constants (not readfrom config) that
TryBuildConnectionStringFromStoredCredentialalways targets. A rig built per the lanerig-setup steps only creates
darlingtest(for the suite) andprobe(for hand-run SQL) - neverdarling-so the old test's "managed" connection attempt failed before it ever reached a settings verdict.
Test plan
dotnet build Darling/Darling.Tests/Darling.Tests.csproj- 0 Warning(s), 0 Error(s).dotnet build Lite.Tests/Lite.Tests.csproj- 0 Warning(s), 0 Error(s).DarlingCliCommandsHostCheckTestslive against the rig (rig-th, port 55983, started in the background,"ready to accept connections" confirmed,
darlingtest/probedropped and recreated first): 7 total, 0failed, 0 skipped.
DarlingMcpStoreHostToolsLiveTests(new, the mcp/viewer least-privilege role proof): 1 total, 0 failed,run twice for stability; confirmed the disposable
host_mcp_test/host_viewer_testroles are dropped aftereach run.
DarlingMcpStoreHostBudgetLiveTests(new, the response-size proof): 1 total, 0 failed.McpToolsListBudgetTests(the tools/list budget pin + layout + total-ceiling checks): 5 total, 0 failed.CrossAppMcpToolInventoryPinTests(Lite.Tests, run once since this branch changed it): 6 total, 0failed.
FleetSweepWebFeedTests+DarlingWebAssetsTests(the web panel's source pins): 59 total, 0 failed.git merge origin/dev- one conflict (McpToolsListBudgetTests.TotalCeilingBytes, both sides had movedit), resolved by summing dev's accumulated raise with this PR's own delta. Kept both change-log comment
blocks.
darlingtest): 13,992 total, 0errors, 4 failed, 29 skipped, 1 not run (765s). Of the 4:
LivePostgresCollectionHygieneTests.EveryClassUsingTheSharedStore_IsSerializedOrDocumentsWhyNot- mine:the new
DarlingMcpStoreHostBudgetLiveTestsreached the shared store without[Collection("live-postgres")]. Fixed (added the attribute) and re-confirmed passing.DarlingMcpFleetSweepToolsTests.TheMirrorsLoggerSeat_TakesTheCapturedServiceLogger_NotAHardcodedNull-mine: a third source pin (beside the two H2a already fixed) still expected the old
BuildReadDispatch(logger)call text. Fixed and re-confirmed passing.TrendPayloadBudgetLiveTests.EveryDefaultAnswer_StaysNearTheBudget_AndTheLargestAnswerStaysUnderTheCap-failed with
"Exception while reading from stream"readingget_file_io_trend, code this PR nevertouches. Read as a transient stream error against the shared rig after a 13,992-test, 765-second run, not
a defect in this diff.
CaptureDownChunkOrderTests.TheShippedRead_ExecutesOnlyTheNewestChunk_AndTheNewestRunDecides_ AgainstDevPostgres- failed on TimescaleDB chunk-ordering logic this PR never touches.5834755849 and 5835678714), ending at
682b465a.git merge origin/devagain, one conflict inMcpToolsListBudgetTests.TotalCeilingBytes, set to the 175,170 bytes the test measured, both change-logcomments kept (
9e3563d1). Live classes on a UTC rig, all 0 failed:DarlingMcpStoreHostToolsLiveTests(1live),
DarlingMcpStoreHostBudgetLiveTests(1 live),DarlingCliCommandsHostCheckTests(7 live),StoreHostProfileCacheTests(4),McpToolsListBudgetTests(5),McpServiceParameterDiSeatCensusTests(1),FleetSweepWebFeedTests(39),DarlingWebAssetsTests(21).TrendPayloadBudgetLiveTestsandCaptureDownChunkOrderTests, the two unexplained failures above, passed.9e3563d1decides.What remains
coordinator's split (part 2a: this PR; part 2b: cloud identity).
Double-check requested (from the earlier draft, still open)
PostgresConfig? postgresConfig = nullthreading throughBuildReadDispatch->MapAllandConfigurePipeline- confirm the optional-parameter, closure-based approach is the right shape rather thanwidening the
ReadToolHandlerdelegate itself.McpStoreHostProfileSeconds = CliStoreReadSeconds * 2 + 5= 25s derivation.sev-Warningis the rightseverity for
stale_after_hardware_change(versussev-Critical- it is a sizing drift, not an outage).CHANGELOG entry
SECTION: Added
ENTRY:
(RAM, CPUs, data volume) and its sizing-relevant settings were invisible from the product; answering "is this
store sized right" needed a remote shell. Added a
get_store_hostMCP/API read reporting the host, thestore's PostgreSQL/TimescaleDB facts, and a verdict per setting (matches, stale after a hardware change,
operator override, or not managed on a bring-your-own store), plus a "Store host" section on the web
dashboard's Fleet Sweeps page showing the same facts and settings table, with stale settings highlighted.
REF:
[Add get_store_host MCP read for the store host profile (#4214) #4282]: Add get_store_host MCP read for the store host profile (#4214) #4282
🤖 Generated with Claude Code
https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ