Repository navigation
get_store_metrics answers with a bounded summary instead of every store object's daily series (#3903) - #3942
Merged
Conversation
…re object's daily series (#3903) The tool returned every store object's latest row AND its daily series on every call. The object count is fixed by the schema (about 250: 72 hypertables, 25 continuous aggregates, 146 background jobs), not by the fleet, so the default answer was ~1.9 MB on every production store and 1,246 KB on DARLING01, more than an MCP client's context. The reads were never the cost; the bytes were. The default is now the store-level blocks plus three ranked lists bounded by limit (default 10): the largest objects, the fastest-growing, and the background jobs (failing in the window first, then closest to cadence). Each row carries its change over the window from delta_since in place of its series. object_kind lists one kind; an exact object_name returns that object's daily series, a partial one lists what contains it. Row shapes split by kind, so job rows no longer carry twenty null byte fields. DARLING01's default answer, measured on its own store_metrics rows: 1,246 KB to 19 KB. The web mirror binds the new parameters, the store alerts' triage page gains a background-jobs section, and the tool description (9,765 to 7,185 characters) and the server-instructions paragraph (3,113 to 1,404) shrink with it (#3898). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif
erikdarlingdata
enabled auto-merge (squash)
September 22, 2026 23:59
erikdarlingdata
added a commit
that referenced
this pull request
Sep 23, 2026
…he first two minutes after midnight UTC (#3963) DailySummary_StopsPaintingPurgedDaysGreen_AndPublishesTheHorizon seeded "today's" run at UtcNow - 2 minutes and then required a day row for today. In a day's first two minutes UTC that seed lands on yesterday, so the test failed CI at 00:01Z on #3942, which is a Darling-only PR. The seed now uses `now` when two minutes back would cross midnight. Test-only; no CHANGELOG entry. Claude-Session: https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
erikdarlingdata
added a commit
that referenced
this pull request
Sep 23, 2026
…3920, #3927, #3931, #3932, #3940, #3942, #3946, #3947, #3950, #3952, #3955, #3956, #3957, #3964, #3965, #3966, #3968, #3972, #3975, #3979, #3980, #3981, #3983, #3984, #3985) (#3989) The wave's fix PRs deliberately carried no CHANGELOG edits (parallel-agent hot-spot protocol); each agent reported its entry and this commit lands them together, byte-verified against origin/dev. Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #3903.
Why
get_store_metricsreturned every store object's latest row and its daily series on every call. The object count is fixed by the schema, not by the fleet: about 250 objects (72 hypertables, 25 continuous aggregates, 146 TimescaleDB jobs, the payload dimensions, named tables and catch-alls). So the default answer was ~1.9 MB on every production store in the issue, about 500k tokens, and evendays_back=1was 347 KB becauseobjects[]alone was 172 KB.The diagnosis is confirmed on DARLING01 (253 objects, 17 days of history). The default response is 1,275,633 bytes:
daily(250 series, 3,518 points)objects(250 rows, 146 of them jobs with ~20 null byte fields each)The SQL is not the cost: on DARLING01's 83,355
store_metricsrows the latest read takes 225 ms and the 30-day daily read 504 ms (EXPLAIN (ANALYZE, BUFFERS), both in-memory quicksorts). The bytes are. The reads are unchanged by this PR.What changes
store(with the whole-store daily growth and the per-server ingest rate),inventory,retention,job_historyandcheckpointer. The per-object series is replaced by three ranked lists. Each list is bounded bylimit(default 10) and carries its match count,truncatedandorder:objects: the largest byte-bearing objects.fastest_growing: ranked bygrowth_bytesover the window.background_jobs: jobs whose failure count grew in the window first, then byduration_vs_cadence_percent.growth_bytes; a job row carriesruns_in_windowandfailures_in_window. Each is measured fromdelta_since(the first daily point in the window) to the last point, the subtraction a caller would make from the series. The value is null, never zero, when there is nothing to state: fewer than two points, a null at either end, or a cumulative counter that went backwards.object_kindlists every object of one kind, bounded bylimit. A kind that is not one of the seven listed kinds is refused (invalid). The store, job_history and checkpointer rows stay blocks and are never rows.object_namedrills in. An exact name (case-insensitive) returns that object's row and its dailyseriesoverdays_back. Otherwise every object whose name contains the text is listed. Exact matching comes first, so a name likewait_statsstill reaches the hypertable alone, not its aggregate and jobs. A filter that matches nothing answersempty./api/read/get_store_metricsadvertises and binds the three new parameters. The store alerts' triage page gains a "Background jobs" section (object_kind=background_job,limit=25). The page's table renders a section's first array, and in the summary the jobs are now a nested list.Behavior changes (wire)
daily.objectsis now a bounded page, not every object.seriesholds one object's daily points when an exactobject_nameis passed, and is null otherwise. The summary'snotesays the per-object series is not in the summary and names the parameter that brings it back.days_backkeeps its meaning and its maximum of 400. It is now also the window for the deltas.get_store_metricson DARLING01 (0 of 3). A saved panel elsewhere that keyed ondailywould needobject_name.get_store_metricsis Darling-only by architecture:CrossAppMcpToolInventoryPinTestsrecords it as such.Test plan
mcp_one): default 1,245.7 KB;days_back=1346.6 KB;days_back=7710.8 KB;days_back=4001,245.7 KB.store_metricsrows copied to the rig: default 19.1 KB, 65 times smaller.days_back=117.9 KB;days_back=718.3 KB;days_back=40019.1 KB; every job (object_kind=background_job,limit=1000) 55.7 KB; one object 12.4 KB; a partial name 13.5 KB.days_back=40045.1 KB (only the whole-store series grows with the window); one object over 400 days 118.7 KB.StoreMetricsSummaryFirstTests(pure): listed kinds, window deltas, exact-first selection, list and growth order, per-view notes, the description's contract together with the code behind it, a source pin that no per-object series projection returns, the web catalog, and the triage section.StoreMetricsSummaryFirstLivePostgresTests, on its own scratch store seeded to production's shape:daily, and the lists' contents checked against arithmetic oracles.delta_sinceandgrowth_bytesagree with the drill-down's own series.days_back=400stays bounded.emptywork.ListedKindsplusSelectObjectsin place of two inequalities.Darling.TestswithDARLING_TEST_PGon the rig, rebased on dev4b6e50c9: 12,499 total, 0 failed, 21 skipped (the gates that needDARLING_TEST_PGRUNTIMEorDARLING_TEST_SQL, and the symlink test without Developer Mode). An earlier run on the pre-rebase build failed onlyPgWaitSamplerLiveTests(31 samples against an upper bound of 30). That test passed in two other full runs and 3 of 3 times alone; it samplespg_stat_activitycluster-wide while the own-store classes run in parallel. Filed as PgWaitSamplerLiveTests can count another test's lock wait: 31 samples from a 30-snapshot cycle under full-suite load #3939.Lite.Testson the pre-rebase build (nothing in Lite or a shared project changed; run because Lite tests read Darling sources): 5,124 total, 0 failed.Follow-up filed
collect.store_metricson every call. At full retention it takes 7.8 s on a CI-sized rig, against 15 ms with a(object_kind, object_name, metric_time DESC)index and a skip scan. That fix needs a migration rung, so it is out of this change.🤖 Generated with Claude Code
https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif