Summary
get_fleet_overview intermittently bands healthy, actively-collecting servers as Offline - no recent collection. The servers are not offline and nothing is failing. query_store's multi-minute runs stretch the per-server sweep interval past the hardcoded 15-minute OfflineThreshold, so the freshness classifier calls a server dark while it is mid-cycle.
This is a display-only defect — the alert engine uses a different, more conservative window and did not fire. But it is actively misleading: it is what sent an investigation chasing five phantom-dark production shards today.
Evidence
Server names are synthetic slugs; the same slug always denotes the same server.
At 20:49Z on 2026-09-02, the use1 store reported 5 offline servers. At 21:07Z, with no intervention, it reported 0.
| server |
20:49Z band |
20:49Z last_collection |
21:07Z band |
21:07Z last_collection |
| shard-alpha |
Offline |
20:31:40 |
Healthy |
21:05:24 |
| shard-beta |
Offline |
20:32:32 |
Healthy |
21:06:17 |
| shard-gamma |
Offline |
20:34:30 |
Healthy |
21:06:06 |
| shard-delta |
Offline |
20:34:20 |
Healthy |
21:07:40 |
| shard-epsilon |
Offline |
20:34:17 |
Warning |
20:58:04 |
Staleness at the 20:49Z snapshot was 15.4–18.2 min for the five. The next-oldest server was 13.9 min stale and banded Critical, not Offline — bracketing OfflineThreshold (15 min) precisely. Nothing else distinguishes the five.
Mechanism
On shard-alpha, over 18:14–21:05, every gap longer than 5 minutes is bounded by a query_store run — 11 of 12 begin the moment query_store completes:
GAP 19.3 min 18:56:19 (query_store, dur=144201ms) -> 19:15:38
GAP 18.9 min 20:31:40 (query_store, dur= 4157ms) -> 20:50:33
GAP 16.5 min 19:53:40 (query_store, dur= 14420ms) -> 20:10:10
GAP 16.3 min 18:36:37 (plan_correction) -> 18:52:57
GAP 15.2 min 19:20:20 (query_store, dur= 56531ms) -> 19:35:33
query_store durations on that server in the window ranged 4,157 ms to 182,222 ms. The sweep does not relaunch promptly after a long query_store cycle, so no collector — including the one-minute ones — records for 6.6–19.3 minutes.
That directly falsifies the premise the band structure rests on, in ServerHealthThresholds:
The fastest scheduled collector's cadence (wait_stats / cpu_utilization / memory_stats etc. all run every minute), so MAX(collection_time) tracks a one-minute rhythm on a healthy server. Freshness bands are multiples of this.
A healthy server here goes 19 minutes. And OfflineThreshold is not a multiple of the cadence — it is a bare 15-minute constant:
public static readonly TimeSpan OfflineThreshold = TimeSpan.FromMinutes(15);
ClassifyFreshness bands purely on age > OfflineThreshold, with no signal for whether a collection is in flight or whether any collector has actually failed.
Blast radius: display only
The self-alert path uses a different and more conservative threshold — DarlingSelfAlertEvaluator.StaleWindow = 30 minutes, configurable via CollectionStaleMinutes. Since the observed gaps peaked at 19.3 min, no alert should have fired, and none did: get_alert_history over 6 hours returns 20 alerts, all deadlocks / CPU / collector-cost regressions, and zero Collection Stopped.
So one condition has two definitions that disagree — the display calls a server dark at 15 minutes while the alert engine has deliberately decided it is not stopped until 30, and only the alert side is configurable.
Prior form
[#2699] fixed this same symptom from a different cause (an unbounded GROUP BY server_id plus a default(DateTime) fallthrough making live servers read as ancient). shard-beta appears in that CHANGELOG's list of affected servers and in today's five. The 48-hour bound and null propagation from #2699 are working correctly — these last_collection values are real, not artifacts. This is a genuine 15-minute gap being misclassified, which #2699 did not address.
Recommended fix — and why it should wait for a measurement
The upstream cause is query_store duration. [#2792] is merged but not yet deployed, and takes the plan fetch from ~60,000 ms CPU to 0.508 s by forcing a hash join where a fixed TVF cardinality estimate was driving QUERY_STORE_PLAN_IN_MEM to be re-executed once per candidate plan_id. That is the thing stretching these sweeps.
Changing OfflineThreshold now and deploying #2792 means we cannot tell which removed the symptom. Deploy #2792 first, then re-measure the sweep gaps. If they no longer cross 15 minutes, no classifier change is warranted.
If gaps still cross 15 minutes after #2792, the principled fix is not to raise a second hardcoded constant but to make the display's Offline decision agree with the alert engine's configurable CollectionStaleMinutes, so one condition has one definition. ServerHealthThresholds exists precisely to stop this class of drift ("Freshness thresholds ... used to live twice at numerically equal but independently-editable values — a drift risk", #1562); the alert engine's window was never folded in.
Separately worth considering regardless: Offline is a strong claim — the enum documents it as "nothing has been collected for long enough to call the server dark" and it renders a red overlay. A server that is merely between sweeps has a band already: Stale.
Summary
get_fleet_overviewintermittently bands healthy, actively-collecting servers asOffline - no recent collection. The servers are not offline and nothing is failing.query_store's multi-minute runs stretch the per-server sweep interval past the hardcoded 15-minuteOfflineThreshold, so the freshness classifier calls a server dark while it is mid-cycle.This is a display-only defect — the alert engine uses a different, more conservative window and did not fire. But it is actively misleading: it is what sent an investigation chasing five phantom-dark production shards today.
Evidence
Server names are synthetic slugs; the same slug always denotes the same server.
At 20:49Z on 2026-09-02, the use1 store reported 5 offline servers. At 21:07Z, with no intervention, it reported 0.
Staleness at the 20:49Z snapshot was 15.4–18.2 min for the five. The next-oldest server was 13.9 min stale and banded
Critical, notOffline— bracketingOfflineThreshold(15 min) precisely. Nothing else distinguishes the five.Mechanism
On
shard-alpha, over 18:14–21:05, every gap longer than 5 minutes is bounded by aquery_storerun — 11 of 12 begin the momentquery_storecompletes:query_storedurations on that server in the window ranged 4,157 ms to 182,222 ms. The sweep does not relaunch promptly after a longquery_storecycle, so no collector — including the one-minute ones — records for 6.6–19.3 minutes.That directly falsifies the premise the band structure rests on, in
ServerHealthThresholds:A healthy server here goes 19 minutes. And
OfflineThresholdis not a multiple of the cadence — it is a bare 15-minute constant:ClassifyFreshnessbands purely onage > OfflineThreshold, with no signal for whether a collection is in flight or whether any collector has actually failed.Blast radius: display only
The self-alert path uses a different and more conservative threshold —
DarlingSelfAlertEvaluator.StaleWindow = 30 minutes, configurable viaCollectionStaleMinutes. Since the observed gaps peaked at 19.3 min, no alert should have fired, and none did:get_alert_historyover 6 hours returns 20 alerts, all deadlocks / CPU / collector-cost regressions, and zero Collection Stopped.So one condition has two definitions that disagree — the display calls a server dark at 15 minutes while the alert engine has deliberately decided it is not stopped until 30, and only the alert side is configurable.
Prior form
[#2699] fixed this same symptom from a different cause (an unbounded
GROUP BY server_idplus adefault(DateTime)fallthrough making live servers read as ancient).shard-betaappears in that CHANGELOG's list of affected servers and in today's five. The 48-hour bound and null propagation from #2699 are working correctly — theselast_collectionvalues are real, not artifacts. This is a genuine 15-minute gap being misclassified, which #2699 did not address.Recommended fix — and why it should wait for a measurement
The upstream cause is
query_storeduration. [#2792] is merged but not yet deployed, and takes the plan fetch from ~60,000 ms CPU to 0.508 s by forcing a hash join where a fixed TVF cardinality estimate was drivingQUERY_STORE_PLAN_IN_MEMto be re-executed once per candidate plan_id. That is the thing stretching these sweeps.Changing
OfflineThresholdnow and deploying #2792 means we cannot tell which removed the symptom. Deploy #2792 first, then re-measure the sweep gaps. If they no longer cross 15 minutes, no classifier change is warranted.If gaps still cross 15 minutes after #2792, the principled fix is not to raise a second hardcoded constant but to make the display's Offline decision agree with the alert engine's configurable
CollectionStaleMinutes, so one condition has one definition.ServerHealthThresholdsexists precisely to stop this class of drift ("Freshness thresholds ... used to live twice at numerically equal but independently-editable values — a drift risk", #1562); the alert engine's window was never folded in.Separately worth considering regardless:
Offlineis a strong claim — the enum documents it as "nothing has been collected for long enough to call the server dark" and it renders a red overlay. A server that is merely between sweeps has a band already:Stale.