Repository navigation
Detect a server-wide wait-stats clear once per pass (#4428) - #4434
Merged
Merged
Conversation
WaitStatsCollector.ReadAsync now peeks the wait_time_ms baselines before any subtraction and calls a pass a clear only when at least 50% of baselined types (with at least 20 baselined) read lower AND the pass's summed wait_time_ms is lower than the summed baselines. On a clear, the three wait_stats delta families are rebased to zero for that server (keeping each baseline's timestamp) through a new ICollectorDeltaCalculator.RebaseFamiliesToZero operation, and an Information line is queued once per server per day through NoteWaitStatsClear/DrainWaitStatsClearWarnings, drained by both hosts beside the existing #3653 A5 discontinuity drain. Refs #4428.
erikdarlingdata
marked this pull request as ready for review
September 26, 2026 16:19
This was referenced Sep 26, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #4428.
Why
On one production store, every monitored server's
wait_statspass shows 100% of wait types reading LOWER at one instant, many times an hour. A scheduled job on those servers clears wait statistics (DBCC SQLPERF('sys.dm_os_wait_stats', CLEAR)); the busiest run every 5 minutes.22% of that store's 24-hour
wait_statsrows had zero interval. The control store had 8%.Not service restarts, and not collection gaps: 42,195 of 42,234 zero-interval instants sat next to a normal pass.
WaitStatsCollector.WritePayloadwas reading each wait type's clear-to-zero as an independent counter reset — the honest "unknowable"(0, 0)pair — for every one of the (typically hundreds of) wait types the clear touched, instead of recognizing the pass as one event and crediting each type's current value since the clear.What changes
WaitStatsCollector.ReadAsyncnow peeks thewait_stats_timefamily's cached baselines (a read-only peek, no mutation) right after the rows are read and before any subtraction happens — the same "decide before any subtraction" point the existing identity-epoch carrier already uses.WaitStatsCollector.DetectClear(a pure, directly-testable rule) calls the pass a clear only when BOTH hold:wait_time_msbaseline > 0, at least 50% read lower now, AND at least 20 such types exist;wait_time_msacross every row is lower than the summed cached baselines.The rule fails toward "not a clear" on any ambiguity — a false positive would rebase (and so silently discard real accrual for) hundreds of wait types that never reset.
On a detected clear,
ICollectorDeltaCalculatorgets a newRebaseFamiliesToZero(serverId, collectorNames)operation (default no-op, like the existingClearServer/ClearGroups/DecideRowmembers), implemented on the sharedCollectorDeltaCalculatorboth hosts run. It zeroes the cached VALUE for the threewait_stats_*delta families on that server, keeping each key's cached timestamp — so the very next ordinary delta call reports "current value minus zero" as the delta since the clear, over the real (kept) interval, instead of reading a shrink as an unknowable reset.A new
PeekBaselines(serverId, collectorName)member (default empty map) lets a caller inspect a family's cached values without updating anything, the same peek-only contractDecideRowalready established for Decide a cached plan's restarted counters row-coherently in query_stats (#4428) #4431.The log line is queued through two new members,
NoteWaitStatsClear(serverId, serverName, nowUtc)andDrainWaitStatsClearWarnings(serverId), throttled to once per server per UTC calendar day (in-memory, like every other per-server throttle on this type). Both hosts (DarlingWorker.cs,Lite/Services/RemoteCollectorService.cs) drain it right beside the existing Brains-review campaign: deferred structural residue (from #3538 / #3539 / #3540 / #3541) #3653 A5 discontinuity drain, at Information:Not changed: a single wait type resetting outside a clear (fewer than 20 baselined types, or a lower majority that doesn't clear the bar) still takes the existing independent per-type path. Server epochs (restart/failover detection) are unaffected. No migration — this is purely an in-memory delta-cache behavior change; no new columns or schema.
Documented limitation, stated in the new
RebaseFamiliesToZerodoc comment: only the slice between the previous pass and the clear is lost — unknowable, since the DMV never reports the pre- and post-clear values in the same row — and the first post-clear pass's per-second rate is therefore slightly understated (its denominator still spans back to the pre-clear baseline's timestamp).#4428builds on#4431's row-coherent reset (DecideRow), already ondev; this PR does not touchquery_statsorQueryStatsCollector.Test plan
New
Darling.Tests.WaitStatsClearDetectionTests(7 facts; the earlier in-test "mutation" fact was removed in favour of the real code mutation below), run in-process on macOS per the repo's xunit v3 recipe:FieldShape_AllLowerAndTotalLower_EveryRowCreditsCurrentValueOverRealInterval: ~900 baselined types, all lower, total lower — runs the SAME rows end to end through the realCollectorDeltaCalculator(seed baselines →RebaseFamiliesToZero→WaitStatsCollector.WritePayload) and asserts every row'sdelta_wait_time_msequals its current value with a real, non-zerosample_interval_seconds. RED on pre-fix code: without the clear detection and rebase, every one of these 900 types independently resets inWritePayload(current value below the stale, un-rebased baseline), sodelta_wait_time_msandsample_interval_secondsare 0 for all 900 rows — the assertionsAssert.All(deltaTimes, d => Assert.Equal(500L, d))andAssert.DoesNotContain(0, intervals)both fail against that shape.SixtyPercentLowerAndTotalLower_IsAClear/FiveOfNineHundredLower_IsNotAClear/FirstPass_NoBaselines_IsNotAClear: boundary shapes for the two-part rule.SixtyPercentLowerButTotalRose_IsNotAClear: the false-positive guard — a majority of types lower but the summed total rose (a few heavy waits absorbing ordinary growth) must NOT be read as a clear.LogLine_OncePerServerPerDay_EvenAcrossTwoClears: twoNoteWaitStatsClearcalls the same UTC day queue one line, drained once; a call the next day queues a fresh one.FieldShape_ThroughObserveWaitStatsClear_EveryRowCreditsCurrentValue: pins the collector's own read step, not just the delta math. Seeds a realCollectorDeltaCalculatorwith one ordinary pass's baselines, then callsWaitStatsCollector.ObserveWaitStatsClear(rows, context)— the new internal methodReadAsyncextracted its post-read block into, with no behaviour change — and asserts it returnstrueand thatWritePayloadon the lower rows afterward credits every row's current value with a real, non-zero interval. There is no manualRebaseFamiliesToZerocall in this test, so aReadAsyncthat stopped calling the detector would leave the assertion failing.dev(a same-shape scenario run in a detached worktree oforigin/dev, which has noObserveWaitStatsClear/rebase step at all — the scenario just seeds baselines and callsWritePayloadwith no rebase):Assert.Equal() Failure: Expected: 500 Actual: 0on all 900 rows, because every wait type reads as its own independent counter reset against the un-rebased baseline.DetectClearchanged fromlowerCount * 2 >= baselinedCountto the unreachablelowerCount * 100 >= baselinedCount * 101, this same pin goes RED:Assert.True() Failure: Expected: True Actual: False(fromAssert.True(WaitStatsCollector.ObserveWaitStatsClear(currentRows, context))). Restoring the bar and rebuilding brings it back GREEN (Darling.Tests.WaitStatsClearDetectionTests: 7 passed, 0 failed). The oldMutation_RequireOverHundredPercentLower_FieldShapeNoLongerDetectsAClearfact, which re-evaluated the rule inside the test instead of mutating the shipped code, is removed.Also run:
QueryStatsRowCoherentResetTests,DeltaSeriesAgeTests,DocCommentHygieneTests(required) — all green alongside the new class (101 passed, 0 failed for that combined run).Lite.Tests.WaitStatsCollectorDefinitionTestswas checked by inspection — it pins the query text and payload column order only, neither of which this change touches — and needs no update.Both
Darling.TestsandLite.Testsbuild 0 warnings / 0 errors with-p:EnableWindowsTargeting=true. Full suite was not run;DocCommentHygieneTestsplus the collector-specific classes above were run explicitly and are green.CHANGELOG entry
SECTION: Fixed
ENTRY:
REF: [Detect a server-wide wait-stats clear once per pass (#4428) #4434]: Detect a server-wide wait-stats clear once per pass (#4428) #4434