Skip to content

get_fleet_overview bands healthy servers Offline when query_store stretches the sweep past 15 minutes #2794

Description

@erikdarlingdata

Summary

get_fleet_overview intermittently bands healthy, actively-collecting servers as Offline - no recent collection. The servers are not offline and nothing is failing. query_store's multi-minute runs stretch the per-server sweep interval past the hardcoded 15-minute OfflineThreshold, so the freshness classifier calls a server dark while it is mid-cycle.

This is a display-only defect — the alert engine uses a different, more conservative window and did not fire. But it is actively misleading: it is what sent an investigation chasing five phantom-dark production shards today.

Evidence

Server names are synthetic slugs; the same slug always denotes the same server.

At 20:49Z on 2026-09-02, the use1 store reported 5 offline servers. At 21:07Z, with no intervention, it reported 0.

server 20:49Z band 20:49Z last_collection 21:07Z band 21:07Z last_collection
shard-alpha Offline 20:31:40 Healthy 21:05:24
shard-beta Offline 20:32:32 Healthy 21:06:17
shard-gamma Offline 20:34:30 Healthy 21:06:06
shard-delta Offline 20:34:20 Healthy 21:07:40
shard-epsilon Offline 20:34:17 Warning 20:58:04

Staleness at the 20:49Z snapshot was 15.4–18.2 min for the five. The next-oldest server was 13.9 min stale and banded Critical, not Offline — bracketing OfflineThreshold (15 min) precisely. Nothing else distinguishes the five.

Mechanism

On shard-alpha, over 18:14–21:05, every gap longer than 5 minutes is bounded by a query_store run — 11 of 12 begin the moment query_store completes:

GAP  19.3 min   18:56:19 (query_store, dur=144201ms) -> 19:15:38
GAP  18.9 min   20:31:40 (query_store, dur=  4157ms) -> 20:50:33
GAP  16.5 min   19:53:40 (query_store, dur= 14420ms) -> 20:10:10
GAP  16.3 min   18:36:37 (plan_correction)           -> 18:52:57
GAP  15.2 min   19:20:20 (query_store, dur= 56531ms) -> 19:35:33

query_store durations on that server in the window ranged 4,157 ms to 182,222 ms. The sweep does not relaunch promptly after a long query_store cycle, so no collector — including the one-minute ones — records for 6.6–19.3 minutes.

That directly falsifies the premise the band structure rests on, in ServerHealthThresholds:

The fastest scheduled collector's cadence (wait_stats / cpu_utilization / memory_stats etc. all run every minute), so MAX(collection_time) tracks a one-minute rhythm on a healthy server. Freshness bands are multiples of this.

A healthy server here goes 19 minutes. And OfflineThreshold is not a multiple of the cadence — it is a bare 15-minute constant:

public static readonly TimeSpan OfflineThreshold = TimeSpan.FromMinutes(15);

ClassifyFreshness bands purely on age > OfflineThreshold, with no signal for whether a collection is in flight or whether any collector has actually failed.

Blast radius: display only

The self-alert path uses a different and more conservative threshold — DarlingSelfAlertEvaluator.StaleWindow = 30 minutes, configurable via CollectionStaleMinutes. Since the observed gaps peaked at 19.3 min, no alert should have fired, and none did: get_alert_history over 6 hours returns 20 alerts, all deadlocks / CPU / collector-cost regressions, and zero Collection Stopped.

So one condition has two definitions that disagree — the display calls a server dark at 15 minutes while the alert engine has deliberately decided it is not stopped until 30, and only the alert side is configurable.

Prior form

[#2699] fixed this same symptom from a different cause (an unbounded GROUP BY server_id plus a default(DateTime) fallthrough making live servers read as ancient). shard-beta appears in that CHANGELOG's list of affected servers and in today's five. The 48-hour bound and null propagation from #2699 are working correctly — these last_collection values are real, not artifacts. This is a genuine 15-minute gap being misclassified, which #2699 did not address.

Recommended fix — and why it should wait for a measurement

The upstream cause is query_store duration. [#2792] is merged but not yet deployed, and takes the plan fetch from ~60,000 ms CPU to 0.508 s by forcing a hash join where a fixed TVF cardinality estimate was driving QUERY_STORE_PLAN_IN_MEM to be re-executed once per candidate plan_id. That is the thing stretching these sweeps.

Changing OfflineThreshold now and deploying #2792 means we cannot tell which removed the symptom. Deploy #2792 first, then re-measure the sweep gaps. If they no longer cross 15 minutes, no classifier change is warranted.

If gaps still cross 15 minutes after #2792, the principled fix is not to raise a second hardcoded constant but to make the display's Offline decision agree with the alert engine's configurable CollectionStaleMinutes, so one condition has one definition. ServerHealthThresholds exists precisely to stop this class of drift ("Freshness thresholds ... used to live twice at numerically equal but independently-editable values — a drift risk", #1562); the alert engine's window was never folded in.

Separately worth considering regardless: Offline is a strong claim — the enum documents it as "nothing has been collected for long enough to call the server dark" and it renders a red overlay. A server that is merely between sweeps has a band already: Stale.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions