Skip to content

Collector Cost Regression tests a 2.0x ratio against a MEAN baseline on collectors whose own p95/mean spread is 2.0-5.0x, so a normal p95 run is a guaranteed firing — #3316 fixed materiality, this is dispersion #3440

Description

@erikdarlingdata

Collector Cost Regression compares one run's ms/run against a mean 14-day baseline at a CostRegressionFactor of 2.0. For a collector whose own p95/mean spread already exceeds 2.0 — which every heavy collector on this fleet does — a perfectly normal p95 run clears the ratio test by construction. The alert then reports a regression that is the upper mode of a known-bimodal distribution.

#3316 fixed the materiality axis (CostRegressionAddedMsFloor = 5000). This is the dispersion axis, and the floor does not screen it, because a slow-but-legitimate run on an expensive collector adds far more than 5 s.

Measured

A live firing, 2026-09-14 23:44:20Z:

index_object_stats, one server
  alert:    17,548.0 ms/run   vs   6,477.2 ms baseline x 2.0   -> fired at 2.7x
  reality:  7 runs in 7 days (once daily)
            avg 8,852 ms      p95 17,935 ms      max 17,935 ms

The run that fired is at the collector's own p95 — 17,548 against a p95 of 17,935 — and the 6,477 ms baseline sits below the collector's own 7-day mean of 8,852. So the alert compared a normal upper-mode run against a baseline lower than the mean and reported 2.7x.

The dispersion is not incidental to this collector. fanout on the same row: 7 databases, of which tempdb alone accounts for 13,718 of 17,935 ms, dominance 5.35. A once-daily fan-out whose cost is dominated by one database's index count is inherently variable run to run.

It generalises — every heavy collector measured exceeds the threshold on its own spread

Same get_collection_health rows, p95_duration_ms / avg_duration_ms:

collector avg ms p95 ms p95/avg
index_object_stats (server A) 8,852 17,935 2.03x
index_object_stats (server B) 6,359 18,754 2.95x
query_store (server A) 4,092 20,626 5.04x
procedure_stats (server A) 1,292 6,123 4.74x

All four exceed CostRegressionFactor = 2.0 on natural variation alone. So for each of them, a p95 run is a guaranteed firing whenever the baseline sits near the mean — no regression required.

And #3316's floor cannot screen these, because the same expense that makes them bimodal makes their excess material: query_store at 5.04x spread adds 16,534 ms on a single p95 run, 3.3x the 5,000 ms floor, from one run.

The product already knows these are two populations

get_collection_health's own description says it, and it is the reason avg, p95 and max are all returned:

Read the three together: avg close to p95 close to max is one population, avg far below p95 is two, and p95 far below max is one pathological run.

peak_cycle_ms is built from p95 rather than the mean for exactly this reason (#2460), and #2701/#2717 applied the same bimodal reasoning to BODY_OVERRUN risk for query_stats and plan_correction. So the codebase reasons about bimodal collector cost on the sweep-pressure surface and ignores it on the alert that pages a human about collector cost. One surface fixed, the sibling carrying the same premise — the #2629 shape.

Why it matters

The failure is not noise volume, it is that the alert cannot discriminate what it claims to detect. A reader cannot tell "this collector got slower" from "this collector had a normal slow day", which is the entire question. #3316's own closing argument was that a ratio finer than the instrument is unfalsifiable at 3 ms; the same argument applies at 9 s when the instrument's own spread is 2x.

A secondary consequence: the baseline is a 14-day mean, so one genuinely pathological run raises it, making the next real regression harder to detect — while a run-of-the-mill p95 day fires. The alert is anti-correlated with what it wants.

What would fix it

In rough order of cost:

  1. Compare against a dispersion-aware baseline. Require the run to exceed baseline_p95 x factor, or mean + k x stddev, rather than mean x factor. The per-collector p95 is already computed and displayed by get_collection_health, so the quantity exists.
  2. Screen on the spread itself — skip the ratio test for any collector whose own p95/avg exceeds the factor, because for those the test cannot discriminate. This is [BUG] Collector Cost Regression's daily-TOTAL floor does not constrain its per-RUN ratio, so a truthful 2x on a 3 ms collector costing 0.16 s/day fires #3316's move (make ineligible what the test cannot measure) applied to the other axis, and it is the smallest change.
  3. Require sustained elevation — N consecutive runs above the bound rather than one. On a once-daily collector that is N days, which is arguably the right cadence for a cost regression as opposed to a slow run.

(2) alone would have suppressed the measured firing, and (1) is the version that keeps detecting real regressions on variable collectors instead of excluding them.

Related: #3316 (the materiality axis, and the "threshold finer than the instrument" argument this extends), #2460 (why peak_cycle_ms uses p95 rather than the mean), #2701 / #2717 (bimodal collector cost recognised on the sweep-pressure surface).

Activity

  1. erikdarlingdata commented on Sep 15, 2026

    @erikdarlingdata
    OwnerAuthor

    Claude posting for Erik Darling

    Reframing this before a lane picks it up, because the issue as filed proposes the wrong class of fix.

    Erik's reaction to the firing was two questions: "the alert seems wrong?" and "what does sending it to an alert channel accomplish?" The first is what this issue describes. The second is the one that should decide the work.

    The arithmetic case, stated more strongly than the issue body does

    The collector runs once per day — 7 runs in 7 days. So its 14-day mean baseline is an average of about 14 single observations, and the quantity being tested is one more single observation from the same distribution. With a measured p95/mean of 2.03x against a CostRegressionFactor of 2.0, the top ~5% of perfectly normal runs clear the threshold by construction. On a daily collector that is roughly 18 firings per server per year with nothing wrong, multiplied across 43 servers and every collector whose spread exceeds the factor (measured: 2.03x, 2.95x, 4.74x, 5.04x).

    A ratio test against a mean cannot discriminate "this collector got slower" from "this collector had a normal slow day" when the distribution's own upper tail exceeds the ratio. That is the same argument #3316 closed with — a threshold finer than the instrument is unfalsifiable — at 9 seconds instead of 3 milliseconds.

    The channel case, which supersedes it

    What action does this alert invite? A once-daily collector cost 11.1 extra seconds today, against a 60,000 ms sweep budget, and it is the monitoring tool's own cost rather than a monitored server's. Nothing is degraded, no data is lost, nothing is at risk, and nothing needs doing before morning. Even a real regression on this metric does not want a page — it wants to be noticed across days, with trend context.

    The history is the argument. #3316 added CostRegressionAddedMsFloor to stop this metric paging on immaterial ratios. This issue proposes adding a dispersion gate to stop it paging on normal variance. Two successive gates bolted on to keep a metric quiet is evidence that the metric is in the wrong delivery channel, not that the third gate will be the one that works. Each gate also narrows what the alert can ever report, so the end state of that path is a paging alert that fires almost never and is therefore also trusted by nobody.

    Collector cost regression is genuinely worth knowing — it is how a collector running 37–100 minutes was found (#2150). The claim is about delivery, not value.

    What I would build instead

    1. Demote the metric from the paging channel to a periodic digest — daily or weekly, one message listing collectors whose cost moved, with avg, p95 and run count beside the ratio so a reader can judge dispersion themselves. That surface can afford to be inclusive precisely because it is not a page, so it needs neither the materiality floor nor a dispersion gate to be useful.
    2. Keep a paging alert only for the shapes that are actually actionable, which are different quantities:
      • abandoned / abandon_rate_pct rising — that is collected data you do not have, which is a real loss and already banded WARNING at 0.5%;
      • a collector's cost rising sustained across many servers at once, which is a deployment or store-side change rather than one server's slow day;
      • peak_cycle_risk reaching BODY_OVERRUN/SATURATED, which means relaunches are being skipped and collection is silently thinning.
    3. If the digest is out of scope for now, the smallest honest change is [BUG] Collector Cost Regression's daily-TOTAL floor does not constrain its per-RUN ratio, so a truthful 2x on a 3 ms collector costing 0.16 s/day fires #3316's own move on this axis: make ineligible what the test cannot measure — skip the ratio for any collector whose own p95/avg exceeds the factor. That suppresses the measured firing and does not pretend the threshold is meaningful where it is not.

    (1)+(2) is the design answer. (3) is the interim if someone wants the noise gone this week.

    One correction to the record

    When Erik asked whether this metric was supposed to be muted, I cited a note that it had been "deliberately left unmuted as the running verification of #3316." He corrected me: that decision was scoped to the PostgreSQL store, and this firing is on a SQL Server store. The reasoning does not transfer, and no decision covers the SQL Server side — the metric has been live there by default since a blanket mute was removed, not by choice. I relayed a store-scoped decision as fleet-wide. Recording it because it is the kind of error that makes a "we already decided this" claim untrustworthy.

  2. erikdarlingdata commented on Sep 15, 2026

    @erikdarlingdata
    OwnerAuthor

    Claude posting for Erik Darling

    #3441 implements the threshold fix and says "Closes #3440", so this issue will close on merge. That is the right disposition for the arithmetic — comparing a single run against a 14-observation mean is the defect, and bounding against the baseline's upper edge is the fix.

    The second question in the comment above — what does sending this to an alert channel accomplish — is not addressed by that PR, so it now lives in #3443 and survives this close. It is scoped explicitly non-blocking on #3441.

  3. erikdarlingdata commented on Sep 15, 2026

    @erikdarlingdata
    OwnerAuthor

    Closed by #3441 (merged to dev).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions