Repository navigation
Collector Cost Regression has needed two gates to stop it paging on unactionable things: the question is delivery, not threshold — a seconds-per-day trend belongs in a digest, not a paging channel #3443
Description
Activity
Claude posting for Erik Darling
Live evidence, measured 00:45:09Z, that makes this issue's case better than the argument above did
Eight
Collector Cost Regressionalerts fired in the same second on the the SQL Server store store. One was delivered.00:45:08 DEV-B plan_correction 3.1x sent=True webhook 00:45:09 PROD-A query_stats 2.4x sent=False undelivered 00:45:09 DEV-B query_stats 2.0x sent=False undelivered 00:45:09 PROD-A plan_cache_stats 3.8x sent=False undelivered 00:45:09 PROD-A ag_replica_states 3.9x sent=False undelivered 00:45:09 PROD-A blocked_process_report 2.2x sent=False undelivered 00:45:09 PROD-A ag_database_replica_states 3.5x sent=False undelivered 00:45:09 PROD-A query_snapshots 2.2x sent=False undeliveredThree things follow, and none of them are threshold problems.
1. The six production alerts were TRUE positives
<prod-server>is genuinely degraded, confirmed independently of this metric:get_collection_log:query_statsran 24-28 s at 23:02-23:04Z and 70-88 s from 23:17Z onward, for an identical 200 rows and ~16 MB — same work, ~3x the time.- The slow phase is
open_ms, notdrain_ms, on collectors that read almost nothing:ag_replica_statesspent 14,231 ms of 14,232 ms in open to return 2 rows / 416 bytes;blocked_process_report9,065 ms for 0 rows;plan_cache_stats5,560 ms for 336 bytes. That is time-to-first-row, not data transfer. get_collector_stall_probeshas 7 probes on this server in 24 h (02:41, 06:55, 12:56, 12:59, 15:58, 22:16, 00:45Z), every one at 0.16-0.26 MB/s — so the slowness is chronic and tonight's 3x is a worsening on top of it.- Every probe rules out the instance:
connect_ms0-15 ms, a second connection ran its DMV query in 6-273 ms,runnable_tasks2-7 againstscheduler_count8,work_queue_length0,pending_disk_io0, and the top waits are idle background waits (DISPATCHER_QUEUE_SEMAPHORE,WAIT_XTP_HOST_WAIT,POPULATE_LOCK_ORDINALS) reported as cumulative-since-startup totals. - The server's health card reads Healthy on every workload axis: CPU 10%, 0 blocking, 0 deadlocks, 453 available threads, 1 thread waiting for CPU.
By the stall-probe tool's own rule — quiet on scheduler, storage and locking points off the instance entirely — this is an off-instance condition in the path to PROD-A. So the metric noticed something real that no other alert on this fleet noticed. This issue is not asking to stop measuring collector cost. It is asking to stop delivering it as eight cards.
2. A server-wide condition fans out into one alert PER COLLECTOR
Six collectors on one server crossed the factor in the same evaluation pass, because the evaluator keys the fire on
(server, collector)and loops. Nothing in the alert says "PROD-A is slow"; six cards each say "one collector on PROD-A is above its baseline." The operator has to reassemble the finding from cards that arrive as peers of a dev box's noise.This is the digest shape's strongest argument. One digest row reading PROD-A: 6 collectors above baseline, 2.2x-3.9x, open-phase dominated is the actual finding. Eight alert cards are a puzzle.
3. The delivery census CANNOT TELL US whether the six reached anyone — and that is #3427
Seven of the eight logged
undeliveredwith a nullsend_error. PerAlertDelivery.FromFanout, that arm means "configured, nothing delivered", and its own remarks state that a webhook post that came back unsuccessful is also reported asundeliveredwith a null error becauseWebhookAlertServicecollapses every per-channel outcome into one bool. Withdelivery.mode = Summaryandper_event_max = 5, a third possibility is that some of the seven were carried inside one summary post whose history row is another alert's.So three materially different outcomes — throttled, failed, or batched-and-delivered — are indistinguishable in the log:
- throttled: the operator correctly saw one card standing for the group;
- failed: six production alerts about a real degradation were silently lost;
- batched: they were delivered inside the dev box's post.
Eight fires exceed
per_event_max = 5, so batching cannot account for all seven regardless. There is no read on this store that resolves which happened. That is worth stating plainly: the store cannot currently answer "did my operator see this?", which is a prerequisite for any argument about whether a channel is working — including the argument in this issue.What this changes about the proposal
Nothing in the shape, but it sharpens the priority order and adds an acceptance test:
- The digest must aggregate per server first, then per collector — a per-collector-only digest would have printed six rows for one condition and reproduced the defect in a quieter channel.
- It should carry the
open_ms/drain_mssplit, because that split is what distinguishes "this query got expensive" from "this server stopped answering promptly", and it is the difference between an app-side finding and an infrastructure one. - Acceptance test, alongside [BUG] 3.4.0 query_store collector runs 37–100 min on Azure SQL DB (3.3.0 median: 4.8 s) — starves all other collectors #2150 and procedure_stats spends 98.1M ms/day rendering plan XML on the monitored servers; capture it on a cadence #2862: whatever ships must render tonight's PROD-A event as ONE legible row identifying the server, and must not let DEV-B's standing regression outrank it.
Also worth recording against the "two gates" argument in the body: the top-of-hour fan-out is paced by the
hasFreshDataPointgate at thecollector_costseries' hourly grain, not bycooldown_minutes(which is 5 on this store). So a third gate on the cooldown would not have prevented any of this either.Closed by #3448 (merged to dev).
- added a commit that references this issue
on Sep 16, 2026
Collector Cost Regressionhas now needed two gates to stop it paging on things nobody can act on — #3316's materiality floor, and #3440/#3441's dispersion-aware baseline. Both are correct fixes to real defects. Neither addresses the question that produced them.The operator's reaction to the firing that opened #3440 was two questions. The first — "the alert seems wrong?" — is #3440, and #3441 answers it. The second is this issue:
The claim
A correct-but-unactionable signal in a paging channel is still the wrong delivery, and collector cost regression is one. Fixing the arithmetic makes it fire less often and fire truthfully; it does not make the resulting message something a human should be interrupted for.
Take the firing at its best — assume #3441 has shipped and the ratio is now genuinely meaningful. The message still reads: a once-daily collector cost 11.1 extra seconds today, on the monitoring tool's own overhead, against a 60,000 ms sweep budget. Nothing is degraded. No data is lost. No monitored server is affected. Nothing needs doing before morning. The right response is to notice it alongside last week's figures, which is a report, not a page.
Two gates and counting is the diagnostic. Each one narrows what the alert may ever report in order to keep it quiet, so the terminal state of that path is an alert that fires almost never — and is therefore trusted by nobody when it does. That is the worst of both: the noise is gone and so is the signal.
This is about delivery, not value
Collector cost regression is worth knowing. It is how a collector running 37–100 minutes was found (#2150), and how a collector spending 98.1M ms/day rendering plan XML was found (#2862). The data should keep being collected and surfaced. The argument is only that a paging channel is the wrong surface for a trend observation whose units are seconds per day.
The existing precedent, and why this differs
The product's answer so far to "this is a standing fact rather than an incident" has been a longer refire interval, same channel:
Stale Mute Rules(#3306) re-states onStaleMuteRefire = 1 dayrather than the shared alert cooldown, deliberately. So this issue proposes a departure and should say why.A stale mute is a standing state that needs a nudge until someone resolves it — the daily repeat is the mechanism, and it stops when the rule goes. A cost regression is a trend observation that needs comparison to be judged at all: the same number is alarming or unremarkable depending on the collector's own dispersion, run count and history, none of which fits in an alert card. Repeating it daily does not help; putting
avg,p95, run count and a week of history beside it does.What I would build
avg/p95/ run count / added-ms-per-day beside the ratio so a reader judges dispersion themselves. Because it is not a page, it can afford to be inclusive: it needs neither [BUG] Collector Cost Regression's daily-TOTAL floor does not constrain its per-RUN ratio, so a truthful 2x on a 3 ms collector costing 0.16 s/day fires #3316's materiality floor nor Measure Collector Cost Regression against the baseline's upper edge, not its mean #3441's dispersion bound to be useful, and it would have surfaced both [BUG] 3.4.0 query_store collector runs 37–100 min on Azure SQL DB (3.3.0 median: 4.8 s) — starves all other collectors #2150 and procedure_stats spends 98.1M ms/day rendering plan XML on the monitored servers; capture it on a cadence #2862 earlier and more legibly than an alert that had to be gated down to stay tolerable.abandoned/abandon_rate_pctrising — an abandoned cycle "stored nothing and advanced no watermark", i.e. collected data you do not have. That is a loss, not a cost, and it already bands WARNING at 0.5%.peak_cycle_riskreachingBODY_OVERRUN/SATURATED— relaunches are being skipped and collection is silently thinning at a multiple of the configured interval.Retention Held,Store Disk Pressure,Compression Job Stuck,Custom Alert Rules UnhealthyandStore Job Over Cadenceare all genuine incidents with actions attached; this issue is not an argument against self-alerts as a category.Explicitly NOT blocking #3441
#3441 fixes a real statistical defect and is strictly better than comparing against a mean — it should ship on its own merits. This issue is the companion, filed separately because #3441 says "Closes #3440", and without a home of its own the delivery question would be archived on a closed issue the moment the arithmetic fix merges.
If the digest is judged not worth building, the honest resolution is to close this saying so — that collector cost stays a page and the gates are the accepted cost of that — rather than leaving the question to resurface with the third gate.
Related: #3316 (first gate), #3440 / #3441 (second gate), #3306 (
StaleMuteRefire, the longer-interval precedent this departs from), #2150 and #2862 (real regressions this metric found, which the digest must not lose), #2701 / #2717 / #2460 (bimodal collector cost already reasoned about on the sweep-pressure surface).