Repository navigation
collection-health reads have no count for the ABANDONED status #2804
Description
Activity
Live evidence, and the practical impact is larger than the title suggests: a collector losing cycles reports
HEALTHYwitherrors: 0.get_collection_healthon a production server, 3-hour window (2026-09-04 00:05Z):procedure_stats status=HEALTHY errors=0 note_count=24 note_summary: "wall-clock budget (120s) reached; cycle abandoned (24 of 1226 runs)" query_stats status=HEALTHY errors=0 note_count=14 note_summary: "wall-clock budget (120s) reached; cycle abandoned (14 of 1226 runs)"38 abandoned cycles in three hours on one server — 2.0% of
procedure_statsruns and 1.1% ofquery_statsruns stored nothing and advanced no watermark. Every one of them surfaces only as free text innote_summary. There is no count field a reader can threshold, alert, or trend on, and the status verdict is unaffected.This matters more now than when it was filed:
ABANDONEDis the status the drain investigation depends on (see #2864), and the abandonment rate is the number that would tell us whether a change helped. Right now it can only be recovered by parsing a sentence.Fixed by #2867, merged to
devasa54abd59. Closing manually:Fixes #Nonly fires on the dev->main release merge, so completed work otherwise sits open until release.What shipped.
abandonedandabandon_rate_pctare now first-class on both MCP tools, with an Abandoned column on the web grid and both WPF grids, and the count feeds the sharedCollectorHealthClassifier, which bands WARNING above a 0.5% rate. Read-side only -- no new column, no schema rung.The threshold came from the fleet. Across 1,639 (server, collector) pairs and 520,455 runs in 24h, only four pairs abandoned anything at all -- 28 runs, 0.005% fleet-wide -- and the per-pair rate distribution is p50 = p75 = p90 = p95 = p99 = 0.000% with a max of 2.157%. 0.5 sits strictly below that observed floor and above the 99th percentile, and keeps a lone abandonment quiet in any window under ~400 runs.
WARNING rather than a new band, because a new string would need learning by four display mappings that fail in opposite directions -- the web's
statusToSevdefaults to "Unknown" while the deprecated Dashboard's converter defaults toTransparent, the same brush it gives HEALTHY. Attribution survives because the abandoned count now sits beside the error count in the same row.One real defect found while wiring it: the fleet rollup builds its own banding row and would have kept calling an abandoning collector HEALTHY while every other surface called it WARNING -- and it compiled, because an unset count defaults to 0. That is the #2779/#2784 shape, so the pin asserts over all four banding reads by name rather than over the one that was broken.
Verified against reality, since the Windows suites cannot run on macOS: the decision table ran red-first against the real built classifier (the four abandonment cases report
HEALTHYwith the threshold disabled), andCollectionHealthSqlwas dumped from the built assembly,PREPAREd verbatim and executed read-only against a production store -- 21 columns,abandoned_countmatching an independentGROUP BYcross-check exactly. Test execution confirmed by arithmetic:Darling.Tests6826 -> 6843 (+17) andLite.Tests3200 -> 3213 (+13), matching the rows added to each.
Follow-up deferred from #2803, which gave a wall-clock-budget-abandoned cycle its own
ABANDONEDstatus.ABANDONEDcurrently lands in no bucket in the four collection-health reads. That is semantically correct — it is neither a success nor an error — and it is why the new status was safe to add without touching any read. But it means the health summary has no count for it, and its message surfaces in neitherlast_error(gatedIN ('ERROR','PERMISSIONS')) norlast_note(gated= 'SUCCESS'). It is visible in the raw Collection Log grid and to anystatus <> 'SUCCESS'query, which is what the originating issue was about, but not in the summary.Adding a parallel
abandoned_countalongside the existingyield_countwould touch roughly 13 sites:ViewerDataService.CollectionHealth.cs(2),Mcp/DarlingDataReader.cs,Lite/Services/LocalDataService.CollectionHealth.csYieldCount = reader.IsDBNull(12) ? 0 : ...YieldCountDarlingMcpDataTools,Lite/Mcp/McpHealthTools)The ordinal readers are the reason this was not bundled: inserting a column shifts every subsequent ordinal, and a mis-shifted ordinal reads a neighbouring column silently rather than failing. That is not a risk worth taking in the same change as the classification fix, particularly with a fleet roll in flight.
Whoever picks this up: add the column at the end of each select list rather than beside
yield_count, so no existing ordinal moves, and consider whether an ordinal-shift guard (asserting each DTO field against its column NAME once at startup) is worth having given four separate readers depend on positional agreement.Also worth deciding at the same time: whether
last_note'sstatus = 'SUCCESS'gate should widen to includeABANDONED, so the budget message reaches the health summary. The existing comment argues against a loose complement because it would dragSESSION_MISSINGandCANCELLEDtext into a column whose claim is that it is not an error — an explicit two-value list would not have that problem, but it is a judgement call about what that column means.