Skip to content

get_store_metrics answers with a bounded summary instead of every store object's daily series (#3903) - #3942

Merged
erikdarlingdata merged 1 commit into
devfrom
feature/3903-store-metrics-payload
Sep 23, 2026
Merged

erikdarlingdata merged 1 commit into
devfrom
feature/3903-store-metrics-payload

Conversation

@erikdarlingdata

Copy link
Copy Markdown
Owner

Closes #3903.

Why

get_store_metrics returned every store object's latest row and its daily series on every call. The object count is fixed by the schema, not by the fleet: about 250 objects (72 hypertables, 25 continuous aggregates, 146 TimescaleDB jobs, the payload dimensions, named tables and catch-alls). So the default answer was ~1.9 MB on every production store in the issue, about 500k tokens, and even days_back=1 was 347 KB because objects[] alone was 172 KB.

The diagnosis is confirmed on DARLING01 (253 objects, 17 days of history). The default response is 1,275,633 bytes:

key bytes share
daily (250 series, 3,518 points) 1,095,762 86%
objects (250 rows, 146 of them jobs with ~20 null byte fields each) 172,348 13.5%
everything else ~7,300 0.6%

The SQL is not the cost: on DARLING01's 83,355 store_metrics rows the latest read takes 225 ms and the 30-day daily read 504 ms (EXPLAIN (ANALYZE, BUFFERS), both in-memory quicksorts). The bytes are. The reads are unchanged by this PR.

What changes

  • Summary first, the default. Every response keeps the store-level blocks: store (with the whole-store daily growth and the per-server ingest rate), inventory, retention, job_history and checkpointer. The per-object series is replaced by three ranked lists. Each list is bounded by limit (default 10) and carries its match count, truncated and order:
    • objects: the largest byte-bearing objects.
    • fastest_growing: ranked by growth_bytes over the window.
    • background_jobs: jobs whose failure count grew in the window first, then by duration_vs_cadence_percent.
  • Each row carries its change over the window in place of its series. A byte-bearing row carries growth_bytes; a job row carries runs_in_window and failures_in_window. Each is measured from delta_since (the first daily point in the window) to the last point, the subtraction a caller would make from the series. The value is null, never zero, when there is nothing to state: fewer than two points, a null at either end, or a cumulative counter that went backwards.
  • object_kind lists every object of one kind, bounded by limit. A kind that is not one of the seven listed kinds is refused (invalid). The store, job_history and checkpointer rows stay blocks and are never rows.
  • object_name drills in. An exact name (case-insensitive) returns that object's row and its daily series over days_back. Otherwise every object whose name contains the text is listed. Exact matching comes first, so a name like wait_stats still reaches the hypertable alone, not its aggregate and jobs. A filter that matches nothing answers empty.
  • Row shapes split by kind. Byte-bearing rows carry sizes, aggregate state and TOAST facts; job rows carry run telemetry. One shape for both had put about 20 null fields on every job row, and jobs are more than half the rows.
  • Web mirror and triage. /api/read/get_store_metrics advertises and binds the three new parameters. The store alerts' triage page gains a "Background jobs" section (object_kind=background_job, limit=25). The page's table renders a section's first array, and in the summary the jobs are now a nested list.
  • The MCP surface costs ~100k tokens of context before the first call: 157 tools, 336 KB of tools/list, 86 KB of instructions #3898 budget. The tool description went from 9,765 to 7,185 characters even with the new contract added. The cut is prose that the response's own notes already state where the value is read: the job-history setting and visibility arms, and the TOAST and checkpointer detail. Every pinned claim stays. The server-instructions paragraph went from 3,113 to 1,404 characters. The four parameter descriptions total 388 characters.

Behavior changes (wire)

  • The default response no longer carries daily. objects is now a bounded page, not every object. series holds one object's daily points when an exact object_name is passed, and is null otherwise. The summary's note says the per-object series is not in the summary and names the parameter that brings it back.
  • Background-job rows no longer carry the byte fields, which were always null on them. Byte-bearing rows no longer carry the job fields, which were always null on them.
  • days_back keeps its meaning and its maximum of 400. It is now also the window for the deltas.
  • No saved custom view reads get_store_metrics on DARLING01 (0 of 3). A saved panel elsewhere that keyed on daily would need object_name.
  • There is no Lite change. get_store_metrics is Darling-only by architecture: CrossAppMcpToolInventoryPinTests records it as such.

Test plan

  • Before, on DARLING01 (the deployed dev build, via mcp_one): default 1,245.7 KB; days_back=1 346.6 KB; days_back=7 710.8 KB; days_back=400 1,245.7 KB.
  • After, on DARLING01's own 83,355 store_metrics rows copied to the rig: default 19.1 KB, 65 times smaller. days_back=1 17.9 KB; days_back=7 18.3 KB; days_back=400 19.1 KB; every job (object_kind=background_job, limit=1000) 55.7 KB; one object 12.4 KB; a partial name 13.5 KB.
  • After, on a production-shaped fixture (253 objects, 43 servers, 400 days): default 18.3 KB; days_back=400 45.1 KB (only the whole-store series grows with the window); one object over 400 days 118.7 KB.
  • New StoreMetricsSummaryFirstTests (pure): listed kinds, window deltas, exact-first selection, list and growth order, per-view notes, the description's contract together with the code behind it, a source pin that no per-object series projection returns, the web catalog, and the triage section.
  • New StoreMetricsSummaryFirstLivePostgresTests, on its own scratch store seeded to production's shape:
    • The default stays under the issue's 50 KB acceptance, with every block present, no daily, and the lists' contents checked against arithmetic oracles.
    • delta_since and growth_bytes agree with the drill-down's own series.
    • days_back=400 stays bounded.
    • Filters, refusals and empty work.
    • The web mirror and the triage page's own section definition bind the new parameters into the same tool.
  • Pins updated deliberately, never loosened:
  • Full Darling.Tests with DARLING_TEST_PG on the rig, rebased on dev 4b6e50c9: 12,499 total, 0 failed, 21 skipped (the gates that need DARLING_TEST_PGRUNTIME or DARLING_TEST_SQL, and the symlink test without Developer Mode). An earlier run on the pre-rebase build failed only PgWaitSamplerLiveTests (31 samples against an upper bound of 30). That test passed in two other full runs and 3 of 3 times alone; it samples pg_stat_activity cluster-wide while the own-store classes run in parallel. Filed as PgWaitSamplerLiveTests can count another test's lock wait: 31 samples from a 30-snapshot cycle under full-suite load #3939.
  • Full Lite.Tests on the pre-rebase build (nothing in Lite or a shared project changed; run because Lite tests read Darling sources): 5,124 total, 0 failed.
  • Mutation checks, each restored:
    • Putting back the unbounded object list fails the source pin and both live tests.
    • Contains-only name matching fails the selection pin and both live tests.
    • Letting a counter reset read as a count fails the window-delta pin.

Follow-up filed

🤖 Generated with Claude Code

https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif

…re object's daily series (#3903)

The tool returned every store object's latest row AND its daily series on
every call. The object count is fixed by the schema (about 250: 72
hypertables, 25 continuous aggregates, 146 background jobs), not by the
fleet, so the default answer was ~1.9 MB on every production store and
1,246 KB on DARLING01, more than an MCP client's context. The reads were
never the cost; the bytes were.

The default is now the store-level blocks plus three ranked lists bounded
by limit (default 10): the largest objects, the fastest-growing, and the
background jobs (failing in the window first, then closest to cadence).
Each row carries its change over the window from delta_since in place of
its series. object_kind lists one kind; an exact object_name returns that
object's daily series, a partial one lists what contains it. Row shapes
split by kind, so job rows no longer carry twenty null byte fields.

DARLING01's default answer, measured on its own store_metrics rows: 1,246 KB
to 19 KB. The web mirror binds the new parameters, the store alerts' triage
page gains a background-jobs section, and the tool description (9,765 to
7,185 characters) and the server-instructions paragraph (3,113 to 1,404)
shrink with it (#3898).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif
@erikdarlingdata
erikdarlingdata enabled auto-merge (squash) September 22, 2026 23:59
@erikdarlingdata
erikdarlingdata merged commit 033cfeb into dev Sep 23, 2026
25 of 28 checks passed
@erikdarlingdata
erikdarlingdata deleted the feature/3903-store-metrics-payload branch September 23, 2026 01:00
erikdarlingdata added a commit that referenced this pull request Sep 23, 2026
…he first two minutes after midnight UTC (#3963)

DailySummary_StopsPaintingPurgedDaysGreen_AndPublishesTheHorizon seeded "today's" run at UtcNow - 2 minutes and then required a day row for today. In a day's first two minutes UTC that seed lands on yesterday, so the test failed CI at 00:01Z on #3942, which is a Darling-only PR. The seed now uses `now` when two minutes back would cross midnight. Test-only; no CHANGELOG entry.


Claude-Session: https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
erikdarlingdata added a commit that referenced this pull request Sep 23, 2026
…3920, #3927, #3931, #3932, #3940, #3942, #3946, #3947, #3950, #3952, #3955, #3956, #3957, #3964, #3965, #3966, #3968, #3972, #3975, #3979, #3980, #3981, #3983, #3984, #3985) (#3989)

The wave's fix PRs deliberately carried no CHANGELOG edits (parallel-agent hot-spot protocol); each agent reported its entry and this commit lands them together, byte-verified against origin/dev.


Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant