Repository navigation
Collector Cost Regression tests a 2.0x ratio against a MEAN baseline on collectors whose own p95/mean spread is 2.0-5.0x, so a normal p95 run is a guaranteed firing — #3316 fixed materiality, this is dispersion #3440
Description
Activity
Claude posting for Erik Darling
Reframing this before a lane picks it up, because the issue as filed proposes the wrong class of fix.
Erik's reaction to the firing was two questions: "the alert seems wrong?" and "what does sending it to an alert channel accomplish?" The first is what this issue describes. The second is the one that should decide the work.
The arithmetic case, stated more strongly than the issue body does
The collector runs once per day — 7 runs in 7 days. So its 14-day mean baseline is an average of about 14 single observations, and the quantity being tested is one more single observation from the same distribution. With a measured
p95/meanof 2.03x against aCostRegressionFactorof 2.0, the top ~5% of perfectly normal runs clear the threshold by construction. On a daily collector that is roughly 18 firings per server per year with nothing wrong, multiplied across 43 servers and every collector whose spread exceeds the factor (measured: 2.03x, 2.95x, 4.74x, 5.04x).A ratio test against a mean cannot discriminate "this collector got slower" from "this collector had a normal slow day" when the distribution's own upper tail exceeds the ratio. That is the same argument #3316 closed with — a threshold finer than the instrument is unfalsifiable — at 9 seconds instead of 3 milliseconds.
The channel case, which supersedes it
What action does this alert invite? A once-daily collector cost 11.1 extra seconds today, against a 60,000 ms sweep budget, and it is the monitoring tool's own cost rather than a monitored server's. Nothing is degraded, no data is lost, nothing is at risk, and nothing needs doing before morning. Even a real regression on this metric does not want a page — it wants to be noticed across days, with trend context.
The history is the argument. #3316 added
CostRegressionAddedMsFloorto stop this metric paging on immaterial ratios. This issue proposes adding a dispersion gate to stop it paging on normal variance. Two successive gates bolted on to keep a metric quiet is evidence that the metric is in the wrong delivery channel, not that the third gate will be the one that works. Each gate also narrows what the alert can ever report, so the end state of that path is a paging alert that fires almost never and is therefore also trusted by nobody.Collector cost regression is genuinely worth knowing — it is how a collector running 37–100 minutes was found (#2150). The claim is about delivery, not value.
What I would build instead
- Demote the metric from the paging channel to a periodic digest — daily or weekly, one message listing collectors whose cost moved, with
avg,p95and run count beside the ratio so a reader can judge dispersion themselves. That surface can afford to be inclusive precisely because it is not a page, so it needs neither the materiality floor nor a dispersion gate to be useful. - Keep a paging alert only for the shapes that are actually actionable, which are different quantities:
abandoned/abandon_rate_pctrising — that is collected data you do not have, which is a real loss and already banded WARNING at 0.5%;- a collector's cost rising sustained across many servers at once, which is a deployment or store-side change rather than one server's slow day;
peak_cycle_riskreachingBODY_OVERRUN/SATURATED, which means relaunches are being skipped and collection is silently thinning.
- If the digest is out of scope for now, the smallest honest change is [BUG] Collector Cost Regression's daily-TOTAL floor does not constrain its per-RUN ratio, so a truthful 2x on a 3 ms collector costing 0.16 s/day fires #3316's own move on this axis: make ineligible what the test cannot measure — skip the ratio for any collector whose own
p95/avgexceeds the factor. That suppresses the measured firing and does not pretend the threshold is meaningful where it is not.
(1)+(2) is the design answer. (3) is the interim if someone wants the noise gone this week.
One correction to the record
When Erik asked whether this metric was supposed to be muted, I cited a note that it had been "deliberately left unmuted as the running verification of #3316." He corrected me: that decision was scoped to the PostgreSQL store, and this firing is on a SQL Server store. The reasoning does not transfer, and no decision covers the SQL Server side — the metric has been live there by default since a blanket mute was removed, not by choice. I relayed a store-scoped decision as fleet-wide. Recording it because it is the kind of error that makes a "we already decided this" claim untrustworthy.
- Demote the metric from the paging channel to a periodic digest — daily or weekly, one message listing collectors whose cost moved, with
Claude posting for Erik Darling
#3441 implements the threshold fix and says "Closes #3440", so this issue will close on merge. That is the right disposition for the arithmetic — comparing a single run against a 14-observation mean is the defect, and bounding against the baseline's upper edge is the fix.
The second question in the comment above — what does sending this to an alert channel accomplish — is not addressed by that PR, so it now lives in #3443 and survives this close. It is scoped explicitly non-blocking on #3441.
Closed by #3441 (merged to dev).
Collector Cost Regressioncompares one run'sms/runagainst a mean 14-day baseline at aCostRegressionFactorof 2.0. For a collector whose own p95/mean spread already exceeds 2.0 — which every heavy collector on this fleet does — a perfectly normal p95 run clears the ratio test by construction. The alert then reports a regression that is the upper mode of a known-bimodal distribution.#3316 fixed the materiality axis (
CostRegressionAddedMsFloor = 5000). This is the dispersion axis, and the floor does not screen it, because a slow-but-legitimate run on an expensive collector adds far more than 5 s.Measured
A live firing, 2026-09-14 23:44:20Z:
The run that fired is at the collector's own p95 — 17,548 against a p95 of 17,935 — and the 6,477 ms baseline sits below the collector's own 7-day mean of 8,852. So the alert compared a normal upper-mode run against a baseline lower than the mean and reported 2.7x.
The dispersion is not incidental to this collector.
fanouton the same row: 7 databases, of whichtempdbalone accounts for 13,718 of 17,935 ms,dominance 5.35. A once-daily fan-out whose cost is dominated by one database's index count is inherently variable run to run.It generalises — every heavy collector measured exceeds the threshold on its own spread
Same
get_collection_healthrows,p95_duration_ms / avg_duration_ms:index_object_stats(server A)index_object_stats(server B)query_store(server A)procedure_stats(server A)All four exceed
CostRegressionFactor = 2.0on natural variation alone. So for each of them, a p95 run is a guaranteed firing whenever the baseline sits near the mean — no regression required.And #3316's floor cannot screen these, because the same expense that makes them bimodal makes their excess material:
query_storeat 5.04x spread adds 16,534 ms on a single p95 run, 3.3x the 5,000 ms floor, from one run.The product already knows these are two populations
get_collection_health's own description says it, and it is the reasonavg,p95andmaxare all returned:peak_cycle_msis built from p95 rather than the mean for exactly this reason (#2460), and #2701/#2717 applied the same bimodal reasoning toBODY_OVERRUNrisk forquery_statsandplan_correction. So the codebase reasons about bimodal collector cost on the sweep-pressure surface and ignores it on the alert that pages a human about collector cost. One surface fixed, the sibling carrying the same premise — the #2629 shape.Why it matters
The failure is not noise volume, it is that the alert cannot discriminate what it claims to detect. A reader cannot tell "this collector got slower" from "this collector had a normal slow day", which is the entire question. #3316's own closing argument was that a ratio finer than the instrument is unfalsifiable at 3 ms; the same argument applies at 9 s when the instrument's own spread is 2x.
A secondary consequence: the baseline is a 14-day mean, so one genuinely pathological run raises it, making the next real regression harder to detect — while a run-of-the-mill p95 day fires. The alert is anti-correlated with what it wants.
What would fix it
In rough order of cost:
baseline_p95 x factor, ormean + k x stddev, rather thanmean x factor. The per-collector p95 is already computed and displayed byget_collection_health, so the quantity exists.p95/avgexceeds the factor, because for those the test cannot discriminate. This is [BUG] Collector Cost Regression's daily-TOTAL floor does not constrain its per-RUN ratio, so a truthful 2x on a 3 ms collector costing 0.16 s/day fires #3316's move (make ineligible what the test cannot measure) applied to the other axis, and it is the smallest change.(2) alone would have suppressed the measured firing, and (1) is the version that keeps detecting real regressions on variable collectors instead of excluding them.
Related: #3316 (the materiality axis, and the "threshold finer than the instrument" argument this extends), #2460 (why
peak_cycle_msuses p95 rather than the mean), #2701 / #2717 (bimodal collector cost recognised on the sweep-pressure surface).