Skip to content

Collector Cost Regression has needed two gates to stop it paging on unactionable things: the question is delivery, not threshold — a seconds-per-day trend belongs in a digest, not a paging channel #3443

Description

@erikdarlingdata

Collector Cost Regression has now needed two gates to stop it paging on things nobody can act on — #3316's materiality floor, and #3440/#3441's dispersion-aware baseline. Both are correct fixes to real defects. Neither addresses the question that produced them.

The operator's reaction to the firing that opened #3440 was two questions. The first — "the alert seems wrong?" — is #3440, and #3441 answers it. The second is this issue:

"what does sending it to an alert channel accomplish?"

The claim

A correct-but-unactionable signal in a paging channel is still the wrong delivery, and collector cost regression is one. Fixing the arithmetic makes it fire less often and fire truthfully; it does not make the resulting message something a human should be interrupted for.

Take the firing at its best — assume #3441 has shipped and the ratio is now genuinely meaningful. The message still reads: a once-daily collector cost 11.1 extra seconds today, on the monitoring tool's own overhead, against a 60,000 ms sweep budget. Nothing is degraded. No data is lost. No monitored server is affected. Nothing needs doing before morning. The right response is to notice it alongside last week's figures, which is a report, not a page.

Two gates and counting is the diagnostic. Each one narrows what the alert may ever report in order to keep it quiet, so the terminal state of that path is an alert that fires almost never — and is therefore trusted by nobody when it does. That is the worst of both: the noise is gone and so is the signal.

This is about delivery, not value

Collector cost regression is worth knowing. It is how a collector running 37–100 minutes was found (#2150), and how a collector spending 98.1M ms/day rendering plan XML was found (#2862). The data should keep being collected and surfaced. The argument is only that a paging channel is the wrong surface for a trend observation whose units are seconds per day.

The existing precedent, and why this differs

The product's answer so far to "this is a standing fact rather than an incident" has been a longer refire interval, same channel: Stale Mute Rules (#3306) re-states on StaleMuteRefire = 1 day rather than the shared alert cooldown, deliberately. So this issue proposes a departure and should say why.

A stale mute is a standing state that needs a nudge until someone resolves it — the daily repeat is the mechanism, and it stops when the rule goes. A cost regression is a trend observation that needs comparison to be judged at all: the same number is alarming or unremarkable depending on the collector's own dispersion, run count and history, none of which fits in an alert card. Repeating it daily does not help; putting avg, p95, run count and a week of history beside it does.

What I would build

  1. A periodic collector-cost digest — daily or weekly, one message, every collector whose per-run cost moved, with avg / p95 / run count / added-ms-per-day beside the ratio so a reader judges dispersion themselves. Because it is not a page, it can afford to be inclusive: it needs neither [BUG] Collector Cost Regression's daily-TOTAL floor does not constrain its per-RUN ratio, so a truthful 2x on a 3 ms collector costing 0.16 s/day fires #3316's materiality floor nor Measure Collector Cost Regression against the baseline's upper edge, not its mean #3441's dispersion bound to be useful, and it would have surfaced both [BUG] 3.4.0 query_store collector runs 37–100 min on Azure SQL DB (3.3.0 median: 4.8 s) — starves all other collectors #2150 and procedure_stats spends 98.1M ms/day rendering plan XML on the monitored servers; capture it on a cadence #2862 earlier and more legibly than an alert that had to be gated down to stay tolerable.
  2. Keep a paging self-alert for the collector-side shapes that genuinely are incidents, which are different quantities:
    • abandoned / abandon_rate_pct rising — an abandoned cycle "stored nothing and advanced no watermark", i.e. collected data you do not have. That is a loss, not a cost, and it already bands WARNING at 0.5%.
    • peak_cycle_risk reaching BODY_OVERRUN / SATURATED — relaunches are being skipped and collection is silently thinning at a multiple of the configured interval.
    • Cost rising sustained across many servers at once — a deployment or store-side change rather than one server's slow day, and the one cost shape that is both real and urgent.
  3. Leave the other self-alerts alone. Retention Held, Store Disk Pressure, Compression Job Stuck, Custom Alert Rules Unhealthy and Store Job Over Cadence are all genuine incidents with actions attached; this issue is not an argument against self-alerts as a category.

Explicitly NOT blocking #3441

#3441 fixes a real statistical defect and is strictly better than comparing against a mean — it should ship on its own merits. This issue is the companion, filed separately because #3441 says "Closes #3440", and without a home of its own the delivery question would be archived on a closed issue the moment the arithmetic fix merges.

If the digest is judged not worth building, the honest resolution is to close this saying so — that collector cost stays a page and the gates are the accepted cost of that — rather than leaving the question to resurface with the third gate.

Related: #3316 (first gate), #3440 / #3441 (second gate), #3306 (StaleMuteRefire, the longer-interval precedent this departs from), #2150 and #2862 (real regressions this metric found, which the digest must not lose), #2701 / #2717 / #2460 (bimodal collector cost already reasoned about on the sweep-pressure surface).

Activity

  1. erikdarlingdata commented on Sep 15, 2026

    @erikdarlingdata
    OwnerAuthor

    Claude posting for Erik Darling

    Live evidence, measured 00:45:09Z, that makes this issue's case better than the argument above did

    Eight Collector Cost Regression alerts fired in the same second on the the SQL Server store store. One was delivered.

    00:45:08  DEV-B  plan_correction             3.1x  sent=True   webhook
    00:45:09  PROD-A   query_stats                 2.4x  sent=False  undelivered
    00:45:09  DEV-B  query_stats                 2.0x  sent=False  undelivered
    00:45:09  PROD-A   plan_cache_stats            3.8x  sent=False  undelivered
    00:45:09  PROD-A   ag_replica_states           3.9x  sent=False  undelivered
    00:45:09  PROD-A   blocked_process_report      2.2x  sent=False  undelivered
    00:45:09  PROD-A   ag_database_replica_states  3.5x  sent=False  undelivered
    00:45:09  PROD-A   query_snapshots             2.2x  sent=False  undelivered
    

    Three things follow, and none of them are threshold problems.

    1. The six production alerts were TRUE positives

    <prod-server> is genuinely degraded, confirmed independently of this metric:

    • get_collection_log: query_stats ran 24-28 s at 23:02-23:04Z and 70-88 s from 23:17Z onward, for an identical 200 rows and ~16 MB — same work, ~3x the time.
    • The slow phase is open_ms, not drain_ms, on collectors that read almost nothing: ag_replica_states spent 14,231 ms of 14,232 ms in open to return 2 rows / 416 bytes; blocked_process_report 9,065 ms for 0 rows; plan_cache_stats 5,560 ms for 336 bytes. That is time-to-first-row, not data transfer.
    • get_collector_stall_probes has 7 probes on this server in 24 h (02:41, 06:55, 12:56, 12:59, 15:58, 22:16, 00:45Z), every one at 0.16-0.26 MB/s — so the slowness is chronic and tonight's 3x is a worsening on top of it.
    • Every probe rules out the instance: connect_ms 0-15 ms, a second connection ran its DMV query in 6-273 ms, runnable_tasks 2-7 against scheduler_count 8, work_queue_length 0, pending_disk_io 0, and the top waits are idle background waits (DISPATCHER_QUEUE_SEMAPHORE, WAIT_XTP_HOST_WAIT, POPULATE_LOCK_ORDINALS) reported as cumulative-since-startup totals.
    • The server's health card reads Healthy on every workload axis: CPU 10%, 0 blocking, 0 deadlocks, 453 available threads, 1 thread waiting for CPU.

    By the stall-probe tool's own rule — quiet on scheduler, storage and locking points off the instance entirely — this is an off-instance condition in the path to PROD-A. So the metric noticed something real that no other alert on this fleet noticed. This issue is not asking to stop measuring collector cost. It is asking to stop delivering it as eight cards.

    2. A server-wide condition fans out into one alert PER COLLECTOR

    Six collectors on one server crossed the factor in the same evaluation pass, because the evaluator keys the fire on (server, collector) and loops. Nothing in the alert says "PROD-A is slow"; six cards each say "one collector on PROD-A is above its baseline." The operator has to reassemble the finding from cards that arrive as peers of a dev box's noise.

    This is the digest shape's strongest argument. One digest row reading PROD-A: 6 collectors above baseline, 2.2x-3.9x, open-phase dominated is the actual finding. Eight alert cards are a puzzle.

    3. The delivery census CANNOT TELL US whether the six reached anyone — and that is #3427

    Seven of the eight logged undelivered with a null send_error. Per AlertDelivery.FromFanout, that arm means "configured, nothing delivered", and its own remarks state that a webhook post that came back unsuccessful is also reported as undelivered with a null error because WebhookAlertService collapses every per-channel outcome into one bool. With delivery.mode = Summary and per_event_max = 5, a third possibility is that some of the seven were carried inside one summary post whose history row is another alert's.

    So three materially different outcomes — throttled, failed, or batched-and-delivered — are indistinguishable in the log:

    • throttled: the operator correctly saw one card standing for the group;
    • failed: six production alerts about a real degradation were silently lost;
    • batched: they were delivered inside the dev box's post.

    Eight fires exceed per_event_max = 5, so batching cannot account for all seven regardless. There is no read on this store that resolves which happened. That is worth stating plainly: the store cannot currently answer "did my operator see this?", which is a prerequisite for any argument about whether a channel is working — including the argument in this issue.

    What this changes about the proposal

    Nothing in the shape, but it sharpens the priority order and adds an acceptance test:

    Also worth recording against the "two gates" argument in the body: the top-of-hour fan-out is paced by the hasFreshDataPoint gate at the collector_cost series' hourly grain, not by cooldown_minutes (which is 5 on this store). So a third gate on the cooldown would not have prevented any of this either.

  2. erikdarlingdata commented on Sep 15, 2026

    @erikdarlingdata
    OwnerAuthor

    Closed by #3448 (merged to dev).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions