Skip to content

Query-snapshots 1s lock-timeout yields (error 1222) are counted as collection errors, so one benign yield paints a Critical health day #1805

Description

@erikdarlingdata

Observed (a production field instance, post-upgrade soak)

Rare, sporadic collector queries hitting the 1-second lock timeout. The timeout itself is the DESIGN WORKING: exactly one collector carries SET LOCK_TIMEOUT 1000 (QuerySnapshotsCollector, both query variants), and yielding after 1s instead of joining a blocking chain is the monitor keeping its never-be-a-blocker promise. The data cost of a hit is near zero: one point-in-time snapshot sweep skipped; the next sweep sees current state; nothing cumulative or watermarked loses anything.

The defect is the CLASSIFICATION, not the yield

No code anywhere handles SQL Server error 1222 specially (verified: zero matches across the collectors, the service, and Common), so the yield surfaces as a generic collection error: an error row in collection_log, counted into collector failure rates — and the daily-health banding treats ONE collection error, alone, as a CRITICAL day (pinned by DailyHealthBandTests.CriticalTriggers_EachAloneIsCritical(collErrors: 1)). Each benign 1-second yield therefore paints a Critical day on the health calendar, which draws repeated operator/agent attention to a non-event. That amplification — not the missed snapshot — is the cost.

Reframe worth keeping

A 1222 on the snapshot collector is evidence about the TARGET's lock contention — signal about the monitored server, currently mislabeled as a monitoring failure.

Fix sketch (classification-only; the collection behavior is correct and stays)

  • Catch error 1222 on the lock-timeout-guarded collector and record it as a deliberate YIELD: its own log shape and counter, distinct from collection errors.
  • Exclude yields from the collection-error count that feeds the Critical band.
  • Sustained clustering (many yields on one server/window) can still surface — as target-contention signal, not monitor failure. Thresholds conservative; defaults over knobs.
  • Do NOT raise the timeout, add hot retries, or extend the guard to other collectors speculatively.

Severity

Low; nothing gates on it. Post-release queue.

Activity

  1. added a commit that references this issue on Jul 28, 2026
  2. erikdarlingdata commented on Jul 28, 2026

    @erikdarlingdata
    OwnerAuthor

    Fixed by #1808 (merged to dev at d2066a2), classification-only per the sketch: a shared YieldsOnLockTimeout flag (true only for query_snapshots) routes both runners' 1222 catches to a new YIELDED collection_log status. The reader audit found every error/success predicate in both apps already exact-match, so yields fall out of the Critical band, collector health, fleet banding, and the collection-stopped self-alert with zero reader changes - and pins now hold those predicates to the exact-match shape so a broadening cannot re-admit them. Yields stay visible as their own counter (Yields column on both Collection Health grids, yields field on both get_collection_health MCP tools) - the target-contention re-frame from this issue's writeup, on the surface where clustering is diagnosable. Deliberately unchanged: the 1s timeout, no retries, no guard extension, and the health-band classifier. DailyHealthBandTests untouched: one REAL collection error alone is still a Critical day.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions