Skip to content

Report an AG failover once, and re-fire a standing replica disconnect (#1696) - #1734

Merged
erikdarlingdata merged 3 commits into
devfrom
feature/1696-ag-dedup-refire
Jul 26, 2026
Merged

erikdarlingdata merged 3 commits into
devfrom
feature/1696-ag-dedup-refire

Conversation

@erikdarlingdata

Copy link
Copy Markdown
Owner

Items 2 and 3 of #1696, in one migration because they ship together. Item 1 (Lite parity) landed in #1726.

Cross-node de-duplication

Every replica in an Availability Group is visible from every node, so a fully-monitored 3-node AG collected the same replica's state three times and reported one failover three times — once per monitored server.

is_local is added to ag_replica_states — from the DMV, not inferred, because matching a monitored server's name against replica_server_name is unreliable across instance names, aliases and listeners — and one server now judges each AG.

Which server is chosen is not arbitrary. A secondary's sys.dm_hadr_* carries only its OWN row (measured on the live fixture), so only the primary sees the whole group. Vantage is ranked None < Remote < Local < LocalPrimary, the best available wins, and a secondary yields to the primary as soon as the primary is monitored. Ties keep the incumbent, so authority cannot oscillate between equally-placed nodes.

The AG edge state drops serverId from its key as part of this. An AG is one object however many of its nodes we watch, so whoever is authoritative reads and writes the same state — which is what lets authority move without re-baselining and losing an alert, or double-firing one.

That also changes what Forget means: it releases the departing server's claim so a survivor can take over, and deliberately keeps the group's edge state, because another node may still be watching it and dropping the state would swallow the next failover.

A NULL is_local reads as "unknown", never as "not local". Rows collected before this migration genuinely do not know, and de-duplicating on a false negative would drop a real alert; a snapshot with no known-local row is treated as un-de-duplicable and every row is kept — the direction that keeps alerts rather than losing them.

Disconnect re-fire

AG Replica Disconnected was a pure edge, so a replica that stayed disconnected for a week announced it once. ag_disconnect_refire_minutes (NOT NULL DEFAULT 0 = off, clamped 0–1440) re-announces on the #1659 pattern:

  • same metric name, because webhook automation keyed on it is exactly what a re-fire exists to re-trigger
  • stamped on delivery only, so an alert suppressed by the master switch cannot consume the window
  • cleared on reconnect, so a later outage starts a fresh episode rather than re-firing late

Default off means nothing starts re-alerting on upgrade.

Migration

Both columns ride V37, claimed after sweeping every worktree on disk and every remote branch — dev was at 36 and nothing claimed 37. Coordinated with ag-topology-builder before touching any shared viewer file.

Two pin tests that needed thought rather than renumbering

  • PgSchemaGeneratorTests: ag_replica_states is now the same two-migration story the database grain already was — V34 created 10 columns, V37 appends is_local, and only their sum equals the generator's output. Reused the existing TruncatedSchema pattern. Its "V34 was not widened in place" sweep is deliberately not copied for this grain: the exact-block assertion above it is strictly stronger, and a substring sweep would false-positive on the database grain's own legitimate is_local column.
  • The two AG tests whose premise this supersedes are rewritten, not renumbered. Per-server scoping and Forget-drops-state are no longer the intended behavior, so they now pin the new contract: a fully-monitored AG reports once, and Forget hands over cleanly with the correct previous role.

The Lite golden DuckDB schema gains is_local too — the generator drives Lite's tables, so a new collector column has to appear there or fresh and upgraded stores diverge.

Testing

Suites green: Darling 3317, Lite 1549, Dashboard 768; Installer builds. No live Installer.Tests categories were run.

New coverage: a fully-monitored 3-node AG reporting two role changes as two alerts rather than six; a primary taking authority from a secondary; the re-fire firing only past its window and only under the same metric name; and the default-off case still being a pure edge after eight hours down.

🤖 Generated with Claude Code

erikdarlingdata and others added 2 commits July 26, 2026 18:53
…#1696)

Items 2 and 3 of #1696, in one migration because they ship together.

CROSS-NODE DE-DUPLICATION. Every replica in an Availability Group is visible
from every node, so a fully-monitored 3-node AG collected the same replica's
state three times and reported one failover THREE times - once per monitored
server. `is_local` is added to ag_replica_states (from the DMV, not inferred:
matching a monitored server's name against replica_server_name is unreliable
across instance names, aliases and listeners), and one server now judges each
AG.

Which server is chosen is not arbitrary. A SECONDARY's sys.dm_hadr_* carries
only its OWN row - measured on the live fixture - so only the primary sees the
whole group. Vantage is therefore ranked None < Remote < Local < LocalPrimary,
the best available wins, and a secondary yields to the primary as soon as the
primary is monitored. Ties keep the incumbent so authority cannot oscillate.

The AG edge state drops serverId from its key as part of this. An AG is one
object however many of its nodes we watch, so whoever is authoritative reads and
writes the SAME state - which is what lets authority move without re-baselining
and losing an alert, or double-firing one. That also changes what Forget means:
it releases the departing server's CLAIM so a survivor can take over, and
deliberately keeps the group's edge state, because another node may still be
watching it and dropping the state would swallow the next failover.

A NULL is_local reads as "unknown", never as "not local". Rows collected before
this migration genuinely do not know, and de-duplicating on a false negative
would drop a real alert; a snapshot with no known-local row is treated as
un-de-duplicable and every row is kept, which is the direction that keeps
alerts rather than losing them.

DISCONNECT RE-FIRE. "AG Replica Disconnected" was a pure edge, so a replica
that stayed disconnected for a week announced it once. ag_disconnect_refire_minutes
(V37, NOT NULL DEFAULT 0 = off, clamped 0-1440) re-announces on the #1659
pattern: same metric name, because webhook automation keyed on it is exactly
what a re-fire exists to re-trigger; stamped on DELIVERY only, so an alert
suppressed by the master switch cannot consume the window; cleared on reconnect
so a later outage starts a fresh episode. Default off means nothing starts
re-alerting on upgrade.

Both columns ride migration V37, claimed after sweeping every worktree on disk
and every remote branch - dev was at 36 and nothing claimed 37.

Two pin tests needed real thought rather than renumbering:

- PgSchemaGeneratorTests: ag_replica_states is now the same two-migration story
  the database grain already was - V34 created 10 columns, V37 appends is_local,
  and only their sum equals the generator's output. Reused the existing
  TruncatedSchema pattern. Its "V34 was not widened in place" sweep is
  deliberately NOT copied for this grain: the exact-block assertion above it is
  strictly stronger, and a substring sweep would false-positive on the DATABASE
  grain's own legitimate is_local column.
- The two AG tests whose premise this supersedes are rewritten rather than
  renumbered, because per-server scoping and Forget-drops-state are no longer
  the intended behavior. They now pin the new contract: a fully-monitored AG
  reports once, and Forget hands over cleanly.

The Lite golden DuckDB schema gains is_local too - the generator drives Lite's
tables, so a new collector column has to appear there or fresh and upgraded
stores diverge.

Suites green: Darling 3290, Lite 1534, Dashboard 768; Installer builds.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
CHANGELOG conflict resolved by keeping BOTH sides. Checked which case it was
first, per ag-topology-builder's warning that "keep both" is not always literal:
the two sides are DIFFERENT entries - my #1734 (AG de-dup + disconnect re-fire)
and their #1731 (Lite AG tab) - not two edits of one shared entry, so a union is
correct here. All three link-refs (#1726, #1731, #1734) survived intact.

No code conflict. Their Lite AG topology reader and my Lite AG alert reader
coexist deliberately - their file documents why the display path is separate
from the alert path - so there is nothing to reconcile there.

Suites green on the merged state: Darling 3317, Lite 1570.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@erikdarlingdata
erikdarlingdata merged commit 862765c into dev Jul 26, 2026
4 checks passed
@erikdarlingdata
erikdarlingdata deleted the feature/1696-ag-dedup-refire branch September 12, 2026 20:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant