Observed (a production field instance, post-upgrade soak)
Rare, sporadic collector queries hitting the 1-second lock timeout. The timeout itself is the DESIGN WORKING: exactly one collector carries SET LOCK_TIMEOUT 1000 (QuerySnapshotsCollector, both query variants), and yielding after 1s instead of joining a blocking chain is the monitor keeping its never-be-a-blocker promise. The data cost of a hit is near zero: one point-in-time snapshot sweep skipped; the next sweep sees current state; nothing cumulative or watermarked loses anything.
The defect is the CLASSIFICATION, not the yield
No code anywhere handles SQL Server error 1222 specially (verified: zero matches across the collectors, the service, and Common), so the yield surfaces as a generic collection error: an error row in collection_log, counted into collector failure rates — and the daily-health banding treats ONE collection error, alone, as a CRITICAL day (pinned by DailyHealthBandTests.CriticalTriggers_EachAloneIsCritical(collErrors: 1)). Each benign 1-second yield therefore paints a Critical day on the health calendar, which draws repeated operator/agent attention to a non-event. That amplification — not the missed snapshot — is the cost.
Reframe worth keeping
A 1222 on the snapshot collector is evidence about the TARGET's lock contention — signal about the monitored server, currently mislabeled as a monitoring failure.
Fix sketch (classification-only; the collection behavior is correct and stays)
- Catch error 1222 on the lock-timeout-guarded collector and record it as a deliberate YIELD: its own log shape and counter, distinct from collection errors.
- Exclude yields from the collection-error count that feeds the Critical band.
- Sustained clustering (many yields on one server/window) can still surface — as target-contention signal, not monitor failure. Thresholds conservative; defaults over knobs.
- Do NOT raise the timeout, add hot retries, or extend the guard to other collectors speculatively.
Severity
Low; nothing gates on it. Post-release queue.
Observed (a production field instance, post-upgrade soak)
Rare, sporadic collector queries hitting the 1-second lock timeout. The timeout itself is the DESIGN WORKING: exactly one collector carries
SET LOCK_TIMEOUT 1000(QuerySnapshotsCollector, both query variants), and yielding after 1s instead of joining a blocking chain is the monitor keeping its never-be-a-blocker promise. The data cost of a hit is near zero: one point-in-time snapshot sweep skipped; the next sweep sees current state; nothing cumulative or watermarked loses anything.The defect is the CLASSIFICATION, not the yield
No code anywhere handles SQL Server error 1222 specially (verified: zero matches across the collectors, the service, and Common), so the yield surfaces as a generic collection error: an error row in collection_log, counted into collector failure rates — and the daily-health banding treats ONE collection error, alone, as a CRITICAL day (pinned by
DailyHealthBandTests.CriticalTriggers_EachAloneIsCritical(collErrors: 1)). Each benign 1-second yield therefore paints a Critical day on the health calendar, which draws repeated operator/agent attention to a non-event. That amplification — not the missed snapshot — is the cost.Reframe worth keeping
A 1222 on the snapshot collector is evidence about the TARGET's lock contention — signal about the monitored server, currently mislabeled as a monitoring failure.
Fix sketch (classification-only; the collection behavior is correct and stays)
Severity
Low; nothing gates on it. Post-release queue.