What happens
A collector cycle abandoned by the #2673 whole-server wall-clock budget is recorded in collection_log with status = 'SUCCESS' — having stored nothing and advanced no watermark.
From the use1 monitoring host's store, over its entire 17-day retention window:
status | n | zero_row
SUCCESS | 36 | 36
Every abandonment ever recorded is a SUCCESS with rows_collected = 0. Two examples on the affected server:
2026-09-02 20:30:29.603220 sql_duration_ms=156925 rows_collected=0 status=SUCCESS
error_message: "wall-clock budget (120s) reached; cycle abandoned"
2026-09-02 22:41:12.627159 sql_duration_ms=122385 rows_collected=0 status=SUCCESS
error_message: "wall-clock budget (120s) reached; cycle abandoned"
Normal procedure_stats runs on that server are ~7–8 s median with 100–150 rows, so these are outliers, not the usual shape.
Why
The abandonment site returns normally rather than throwing:
return new CollectorRunResult(0, sqlSlice.ElapsedMilliseconds, 0, $"wall-clock budget ({budgetSeconds}s) reached; cycle abandoned");
so it lands on the ordinary success path, whose status is a hardcoded literal in both hosts — DarlingWorker and Lite's RunCollectorAsync. It travels through the Note / telemetry.Note channel, which is documented as being for "a successful-but-empty run worth explaining" (#1837) and explicitly does not change status. #2673 reused that channel for a case that is not a successful empty run.
Why it matters more than a wrong label
- It claims a collection that did not happen.
DarlingSelfAlertEvaluator.ReadCollectionSignalsAsync takes both last_success and recent_success from status IN ('SUCCESS', 'SKIPPED'). A collector abandoning every cycle therefore reads as perpetually fresh.
- It is invisible to the obvious check. An hourly health check counting non-SUCCESS
collection_log rows, or grepping the app log for cycle abandoned, reports zero while this is happening — which is what happened repeatedly on 2026-09-02.
- It pollutes the note channel.
last_note / note_count are gated on status = 'SUCCESS', whose whole claim is that the run succeeded.
Same family as the cancelled watermark read returning null (indistinguishable from a first run) and the ON CONFLICT batch loss invisible to collection_log: something reported clean while collecting nothing.
Is the underlying slowness new? No.
Worth stating because the fix makes these newly visible. procedure_stats daily duration on the affected server across the full retention window (over_60s = runs above 60 s):
| day |
runs |
p50 |
p95 |
max |
over_60s |
| 08-17 |
32 |
8.3 s |
178.3 s |
246.4 s |
6 |
| 08-18 |
81 |
7.3 s |
98.6 s |
265.4 s |
5 |
| 08-19 |
129 |
7.4 s |
90.6 s |
278.1 s |
8 |
| 08-22 |
143 |
6.7 s |
75.6 s |
221.8 s |
9 |
| 08-27 |
106 |
7.6 s |
18.0 s |
176.6 s |
2 |
| 09-01 |
127 |
7.6 s |
114.6 s |
125.0 s |
10 |
| 09-02 |
126 |
8.4 s |
120.3 s |
156.9 s |
11 |
Slow runs happened every day of the window, and the worst maxima (246–278 s) are from before #2673 landed on 2026-08-28. The budget did not create the slowness; it capped it — which is why max falls from ~250–278 s to ~121–157 s after 08-29.
Abandonment counts rise 4 → 3 → 4 → 10 → 15 per day, but the first abandonment is 08-29 and #2673 landed 08-28, so there is no pre-feature baseline and that rise is not evidence of a regression. p95 is noisy across the whole window (15 s to 178 s), and 114/120 s on 09-01/02 sits inside the range already seen on 08-17 (178 s) and 08-19 (90 s).
Fleet control — procedure_stats over-60 s runs per day with the affected server excluded — is 0–13/day across ~5,000 daily runs with no trend, and 1 on 09-02. This is an the affected server-specific, long-standing characteristic.
So no separate performance issue is filed. Filing "abandonment is rising" as a regression would be a false trend claim off an instrument installed mid-window.
Concentration (whole window): the affected server procedure_stats 22, the affected server query_stats 13, a second server procedure_stats 1. Note query_stats abandons too, not just procedure_stats.
Fix
Give the abandoned cycle its own status. Detail and the rejected alternatives are in the PR.
What happens
A collector cycle abandoned by the #2673 whole-server wall-clock budget is recorded in
collection_logwithstatus = 'SUCCESS'— having stored nothing and advanced no watermark.From the use1 monitoring host's store, over its entire 17-day retention window:
Every abandonment ever recorded is a SUCCESS with
rows_collected = 0. Two examples onthe affected server:Normal
procedure_statsruns on that server are ~7–8 s median with 100–150 rows, so these are outliers, not the usual shape.Why
The abandonment site returns normally rather than throwing:
so it lands on the ordinary success path, whose status is a hardcoded literal in both hosts —
DarlingWorkerand Lite'sRunCollectorAsync. It travels through theNote/telemetry.Notechannel, which is documented as being for "a successful-but-empty run worth explaining" (#1837) and explicitly does not change status. #2673 reused that channel for a case that is not a successful empty run.Why it matters more than a wrong label
DarlingSelfAlertEvaluator.ReadCollectionSignalsAsynctakes bothlast_successandrecent_successfromstatus IN ('SUCCESS', 'SKIPPED'). A collector abandoning every cycle therefore reads as perpetually fresh.collection_logrows, or grepping the app log forcycle abandoned, reports zero while this is happening — which is what happened repeatedly on 2026-09-02.last_note/note_countare gated onstatus = 'SUCCESS', whose whole claim is that the run succeeded.Same family as the cancelled watermark read returning
null(indistinguishable from a first run) and theON CONFLICTbatch loss invisible tocollection_log: something reported clean while collecting nothing.Is the underlying slowness new? No.
Worth stating because the fix makes these newly visible.
procedure_statsdaily duration on the affected server across the full retention window (over_60s= runs above 60 s):Slow runs happened every day of the window, and the worst maxima (246–278 s) are from before #2673 landed on 2026-08-28. The budget did not create the slowness; it capped it — which is why
maxfalls from ~250–278 s to ~121–157 s after 08-29.Abandonment counts rise 4 → 3 → 4 → 10 → 15 per day, but the first abandonment is 08-29 and #2673 landed 08-28, so there is no pre-feature baseline and that rise is not evidence of a regression. p95 is noisy across the whole window (15 s to 178 s), and 114/120 s on 09-01/02 sits inside the range already seen on 08-17 (178 s) and 08-19 (90 s).
Fleet control —
procedure_statsover-60 s runs per day with the affected server excluded — is 0–13/day across ~5,000 daily runs with no trend, and 1 on 09-02. This is an the affected server-specific, long-standing characteristic.So no separate performance issue is filed. Filing "abandonment is rising" as a regression would be a false trend claim off an instrument installed mid-window.
Concentration (whole window): the affected server
procedure_stats22, the affected serverquery_stats13, a second serverprocedure_stats1. Notequery_statsabandons too, not justprocedure_stats.Fix
Give the abandoned cycle its own status. Detail and the rejected alternatives are in the PR.