Skip to content

A wall-clock-budget-abandoned cycle is recorded as SUCCESS in collection_log #2801

Description

@erikdarlingdata

What happens

A collector cycle abandoned by the #2673 whole-server wall-clock budget is recorded in collection_log with status = 'SUCCESS' — having stored nothing and advanced no watermark.

From the use1 monitoring host's store, over its entire 17-day retention window:

status  |  n | zero_row
SUCCESS | 36 |       36

Every abandonment ever recorded is a SUCCESS with rows_collected = 0. Two examples on the affected server:

2026-09-02 20:30:29.603220  sql_duration_ms=156925  rows_collected=0  status=SUCCESS
    error_message: "wall-clock budget (120s) reached; cycle abandoned"
2026-09-02 22:41:12.627159  sql_duration_ms=122385  rows_collected=0  status=SUCCESS
    error_message: "wall-clock budget (120s) reached; cycle abandoned"

Normal procedure_stats runs on that server are ~7–8 s median with 100–150 rows, so these are outliers, not the usual shape.

Why

The abandonment site returns normally rather than throwing:

return new CollectorRunResult(0, sqlSlice.ElapsedMilliseconds, 0, $"wall-clock budget ({budgetSeconds}s) reached; cycle abandoned");

so it lands on the ordinary success path, whose status is a hardcoded literal in both hosts — DarlingWorker and Lite's RunCollectorAsync. It travels through the Note / telemetry.Note channel, which is documented as being for "a successful-but-empty run worth explaining" (#1837) and explicitly does not change status. #2673 reused that channel for a case that is not a successful empty run.

Why it matters more than a wrong label

  1. It claims a collection that did not happen. DarlingSelfAlertEvaluator.ReadCollectionSignalsAsync takes both last_success and recent_success from status IN ('SUCCESS', 'SKIPPED'). A collector abandoning every cycle therefore reads as perpetually fresh.
  2. It is invisible to the obvious check. An hourly health check counting non-SUCCESS collection_log rows, or grepping the app log for cycle abandoned, reports zero while this is happening — which is what happened repeatedly on 2026-09-02.
  3. It pollutes the note channel. last_note / note_count are gated on status = 'SUCCESS', whose whole claim is that the run succeeded.

Same family as the cancelled watermark read returning null (indistinguishable from a first run) and the ON CONFLICT batch loss invisible to collection_log: something reported clean while collecting nothing.

Is the underlying slowness new? No.

Worth stating because the fix makes these newly visible. procedure_stats daily duration on the affected server across the full retention window (over_60s = runs above 60 s):

day runs p50 p95 max over_60s
08-17 32 8.3 s 178.3 s 246.4 s 6
08-18 81 7.3 s 98.6 s 265.4 s 5
08-19 129 7.4 s 90.6 s 278.1 s 8
08-22 143 6.7 s 75.6 s 221.8 s 9
08-27 106 7.6 s 18.0 s 176.6 s 2
09-01 127 7.6 s 114.6 s 125.0 s 10
09-02 126 8.4 s 120.3 s 156.9 s 11

Slow runs happened every day of the window, and the worst maxima (246–278 s) are from before #2673 landed on 2026-08-28. The budget did not create the slowness; it capped it — which is why max falls from ~250–278 s to ~121–157 s after 08-29.

Abandonment counts rise 4 → 3 → 4 → 10 → 15 per day, but the first abandonment is 08-29 and #2673 landed 08-28, so there is no pre-feature baseline and that rise is not evidence of a regression. p95 is noisy across the whole window (15 s to 178 s), and 114/120 s on 09-01/02 sits inside the range already seen on 08-17 (178 s) and 08-19 (90 s).

Fleet control — procedure_stats over-60 s runs per day with the affected server excluded — is 0–13/day across ~5,000 daily runs with no trend, and 1 on 09-02. This is an the affected server-specific, long-standing characteristic.

So no separate performance issue is filed. Filing "abandonment is rising" as a regression would be a false trend claim off an instrument installed mid-window.

Concentration (whole window): the affected server procedure_stats 22, the affected server query_stats 13, a second server procedure_stats 1. Note query_stats abandons too, not just procedure_stats.

Fix

Give the abandoned cycle its own status. Detail and the rejected alternatives are in the PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions