Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

### Added

- **Darling: an AG failover is reported once, not once per monitored node, and a standing replica disconnect can re-alert** ([#1734]) - completes #1696. **Cross-node de-duplication:** every replica in an Availability Group is visible from every node, so a fully-monitored 3-node AG collected the same replica three times and reported one failover THREE times. `ag_replica_states` gains `is_local` (read from the DMV rather than inferred - matching a monitored server against `replica_server_name` is unreliable across instance names, aliases and listeners) and one server now judges each AG. Which one is not arbitrary: a SECONDARY's `sys.dm_hadr_*` carries only its OWN row, measured on the live fixture, so only the primary sees the whole group - vantage is ranked None < Remote < Local < LocalPrimary, the best available wins, a secondary yields to the primary as soon as the primary is monitored, and ties keep the incumbent so authority cannot oscillate. The AG edge state drops the server id from its key as part of this, because an AG is ONE object however many of its nodes are watched; that is what lets authority move without re-baselining and losing an alert, or double-firing one. Removing a server now releases its CLAIM rather than dropping the group's state, since another node may still be watching it. A NULL `is_local` reads as "unknown" and never as "not local" - rows predating the migration genuinely do not know, and de-duplicating on a false negative would drop a real alert. **Disconnect re-fire:** "AG Replica Disconnected" was a pure edge, so a replica disconnected for a week announced it once; `ag_disconnect_refire_minutes` (default 0 = off, clamped 0-1440) re-announces on the [#1674] pattern - same metric name so webhook automation keyed on it re-triggers, stamped on DELIVERY only so a suppressed alert cannot consume the window, and cleared on reconnect. Both columns ride store migration V37.

- **Availability Groups tab in Lite, and one shared AG projection behind every surface** ([#1731]) - Lite gets the AG topology tab Darling's viewer got in #1722, reading its own local DuckDB, and the banding and card-shaping rules move into `PerformanceMonitor.Common.AgTopology` so the web dashboard, the Darling viewer and Lite all draw the same AG from one set of rules. There were two copies after #1722; there is now one, and twin test matrices in both test projects fail together if the shared rules change - which is what makes "these surfaces cannot disagree" a checked claim instead of an intention. The brushes stay per app (Common cannot reference WPF): the model carries the verdict, a `SeverityBrushConverter` carries the colour, which also removed the brush properties from the models entirely. **Lite needed its own read, and that is the part worth knowing.** The AG alert read from #1696 looks reusable and is not: it selects five columns where a topology view needs twelve, and it DROPS rows whose `ag_name` or `replica_server_name` is NULL - correctly, because those are the alert's state-key identity and a row that cannot be keyed cannot have an edge tracked for it. A topology view must do the opposite, since those NULLs appear under WSFC quorum loss, which is exactly when an operator opens the page. Lite's read also correlates its latest-snapshot MAX per server rather than globally, so a lagging instance is not erased by a livelier one's newer timestamp. Same hidden-until-AG-rows reveal as Darling's, same honest empty state, same worst-first ordering. Pointed at a secondary, Lite shows one replica and no primary - that is what that server can actually see, and saying so beats implying it knows the whole group.
- **Lite gets the Availability Group alerts, off a shared policy** ([#1726]) - Lite has collected both AG grains since [#1688] but could not tell you a replica had failed over, disconnected, fallen behind, or had data movement suspended; only Darling could. Lite now raises the same four conditions under the SAME metric names, so a webhook keyed on `AG Failover` matches whichever app sent it. The rules are shared rather than reimplemented: a new `PerformanceMonitor.Common.AgAlertPolicy` holds the readings, the metric-name consts and every pure decision, on the `ConnectionAlertPolicy` pattern - each app owns only its own edge STATE and its own delivery, which is the part that is genuinely app-shaped. Darling was refactored onto it with no behavior change (its whole suite passes untouched, which is why the lift was done as its own step). Lite reads the latest snapshot of each grain from DuckDB, each freshness-gated on its OWN collection time because the two AG collectors are scheduled independently and a stale database-grain snapshot must not be vouched for by a healthy replica-grain one; the evaluator is WPF-free so it pins directly, and delivery goes through the same mute-check and send path every other Lite alert uses, inheriting muting, silencing, the combined history row and the email/webhook fan-out. Three settings mirror Darling's V35 knobs (`notify_ag_health` on, `ag_lag_alert_seconds` 300, `ag_redo_queue_alert_kb` 0 = off), clamped to the same ranges on load AND on save so the stored value and the effective value cannot disagree. Server removal drops the AG state, or a remove-then-re-add would compare the new first sighting against the old role and page a phantom failover. Every rule the earlier AG work paid for carries over: a first sighting is a silent baseline, NULL is never a transition, and a suspended row may raise an alarm but may never clear one - including that a suspended secondary drifting past the threshold still fires. Lite's collector-coverage ratchet went red on this exactly as designed - both AG tables were allow-listed as collect-only, and adding a reader forced the entries out - so that allow-list is now empty.

Expand Down Expand Up @@ -1717,6 +1719,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
[#1716]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/1716
[#1710]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/1710
[#1726]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/1726
[#1734]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/1734
[#1725]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/1725
[#1730]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/1730
[#1727]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/1727
Expand Down
4 changes: 2 additions & 2 deletions Darling/Darling.Tests/DarlingObservabilityTests.cs
Original file line number Diff line number Diff line change
Expand Up @@ -76,8 +76,8 @@ public void MigrationScripts_AreRegisteredInAscendingOrder_V34AgCollectors_V36Ag
Assert.Equal(33, PgMigrations.Scripts[32].Version);
/* The newest migration is asserted by identity rather than by ordinal: this ladder is walked by every
stacked branch at once, and a positional pin turns each addition into a conflict for the next. */
Assert.Equal(36, PgMigrations.Scripts[^1].Version);
Assert.Equal(36, StorageVersion.SchemaVersion);
Assert.Equal(37, PgMigrations.Scripts[^1].Version);
Assert.Equal(37, StorageVersion.SchemaVersion);

/* V34 (#991) creates the two Availability Group collector tables. Schema-qualified collect.* and
CREATE TABLE IF NOT EXISTS, per the file's additive-create idiom (V29): a no-op on a fresh store
Expand Down
189 changes: 155 additions & 34 deletions Darling/Darling.Tests/DarlingSelfAlertTests.cs
Original file line number Diff line number Diff line change
Expand Up @@ -152,6 +152,9 @@ private sealed class Harness
public int AgLagAlertSeconds { get; set; } = 300;
public long AgRedoQueueAlertKb { get; set; }

/// <summary>#1696 (V37): AG disconnect re-fire minutes. Default 0 = off, the shipped behavior.</summary>
public int AgDisconnectRefireMinutes { get; set; }

public DateTime Now { get; set; } = new(2026, 7, 1, 12, 0, 0, DateTimeKind.Utc);

/// <summary>#1681: captures what the evaluator writes to the service log, so the firing/recovery pair
Expand All @@ -167,7 +170,8 @@ private sealed class Harness
connectionRefireMinutes: () => ConnectionRefireMinutes,
notifyAgHealth: () => NotifyAgHealth,
agLagAlertSeconds: () => AgLagAlertSeconds,
agRedoQueueAlertKb: () => AgRedoQueueAlertKb);
agRedoQueueAlertKb: () => AgRedoQueueAlertKb,
agDisconnectRefireMinutes: () => AgDisconnectRefireMinutes);
}

/* ---------------- #991 Availability Group fixtures ---------------- */
Expand All @@ -177,8 +181,9 @@ private sealed class Harness
private const string Db = "Sales";

private static AgReplicaReading ReplicaRow(
string? role = "SECONDARY", string? connected = "CONNECTED", string replica = Replica, string ag = Ag) =>
new(ag, replica, role, connected);
string? role = "SECONDARY", string? connected = "CONNECTED", string replica = Replica, string ag = Ag,
bool? isLocal = null) =>
new(ag, replica, role, connected, isLocal);

private static AgDatabaseReading DatabaseRow(
long? lagSeconds = 0,
Expand Down Expand Up @@ -1443,28 +1448,28 @@ public async Task AgSyncFellBehind_BothTriggersOff_NeitherFiresNorResolves()
}

[Fact]
public async Task AgSyncFellBehind_RecoverySweep_IsScopedToTheServerBeingEvaluated()
public async Task AgSyncFellBehind_OneAgIsJudgedByOneServer_SoAFullyMonitoredAgReportsItOnce()
{
var h = new Harness();
var e = h.Build();

const int otherServerId = 515151;
const int nodeA = 100001;
const int nodeB = 100002;

/* Two servers each have a lagging database. */
await e.ApplyAgDatabaseHealthAsync(ServerId, Name, new[] { DatabaseRow(lagSeconds: 600) }, Ct);
await e.ApplyAgDatabaseHealthAsync(otherServerId, "OTHER-SRV", new[] { DatabaseRow(lagSeconds: 600) }, Ct);
Assert.Equal(2, h.Deliverer.Outcomes.Count);
/* #1696: both nodes are monitored and BOTH see the same AG database, because every replica is
visible from every node. Only one may judge it, or a 3-node AG reports one problem three times. */
await e.ApplyAgDatabaseHealthAsync(nodeA, "NODE-A", new[] { DatabaseRow(lagSeconds: 600) }, Ct);
await e.ApplyAgDatabaseHealthAsync(nodeB, "NODE-B", new[] { DatabaseRow(lagSeconds: 600) }, Ct);

/* The FIRST server catches up. Its snapshot says nothing about the second server, so a recovery sweep
that walked every tracked key would silently resolve the other server's live alert. */
await e.ApplyAgDatabaseHealthAsync(ServerId, Name, new[] { DatabaseRow(lagSeconds: 0) }, Ct);
var resolution = Assert.Single(h.History.Records);
Assert.Equal(Key, resolution.ServerId);
var fired = Assert.Single(h.Deliverer.Outcomes);
Assert.Equal("AG Sync Fell Behind", fired.MetricName);

/* Proof the other server's state survived: it is still inside its cooldown, so it stays quiet rather
than re-firing as a fresh episode. */
await e.ApplyAgDatabaseHealthAsync(otherServerId, "OTHER-SRV", new[] { DatabaseRow(lagSeconds: 600) }, Ct);
Assert.Equal(2, h.Deliverer.Outcomes.Count);
/* The recovery is announced once too, by the same authoritative node. */
await e.ApplyAgDatabaseHealthAsync(nodeA, "NODE-A", new[] { DatabaseRow(lagSeconds: 0) }, Ct);
await e.ApplyAgDatabaseHealthAsync(nodeB, "NODE-B", new[] { DatabaseRow(lagSeconds: 0) }, Ct);

var resolution = Assert.Single(h.History.Records);
Assert.Equal("AG Sync Recovered", resolution.MetricName);
}

/* ---------------- #991: database suspended ---------------- */
Expand Down Expand Up @@ -1521,6 +1526,119 @@ public async Task AgDatabaseSuspended_AlreadySuspendedAtFirstSighting_IsASilentB
Assert.Equal("no reason reported", fired.CurrentValue);
}

/* ---------------- #1696: fleet de-dup + disconnect re-fire ---------------- */

[Fact]
public async Task AgFailover_FullyMonitoredThreeNodeAg_ReportsTheFailoverOnce()
{
var h = new Harness();
var e = h.Build();

/* Every replica is visible from EVERY node, so all three monitored servers report the same two
replica rows. Before #1696 that meant one failover paged three times. */
var beforeFailover = new[]
{
ReplicaRow(role: "PRIMARY", replica: "NODE1"),
ReplicaRow(role: "SECONDARY", replica: "NODE2"),
};
var afterFailover = new[]
{
ReplicaRow(role: "SECONDARY", replica: "NODE1"),
ReplicaRow(role: "PRIMARY", replica: "NODE2"),
};

foreach (var serverId in new[] { 100001, 100002, 100003 })
{
await e.ApplyAgReplicaHealthAsync(serverId, "NODE" + serverId, beforeFailover, Ct);
}

Assert.Empty(h.Deliverer.Outcomes);

foreach (var serverId in new[] { 100001, 100002, 100003 })
{
await e.ApplyAgReplicaHealthAsync(serverId, "NODE" + serverId, afterFailover, Ct);
}

/* Two replicas changed role, so two alerts — NOT six. */
Assert.Equal(2, h.Deliverer.Outcomes.Count);
Assert.All(h.Deliverer.Outcomes, o => Assert.Equal("AG Failover", o.MetricName));
}

[Fact]
public async Task AgAuthority_PrimaryTakesOverFromASecondary_BecauseOnlyThePrimarySeesTheWholeGroup()
{
var h = new Harness();
var e = h.Build();

/* A monitored SECONDARY claims the AG first — its sys.dm_hadr_* is a one-row self-view. */
await e.ApplyAgReplicaHealthAsync(
200001, "SEC", new[] { ReplicaRow(role: "SECONDARY", replica: "NODE2", isLocal: true) }, Ct);

/* The PRIMARY is then monitored. Its vantage is strictly better, so it takes over and its view of
NODE1 is judged — which the secondary could never have supplied. */
await e.ApplyAgReplicaHealthAsync(200002, "PRI", new[]
{
ReplicaRow(role: "PRIMARY", replica: "NODE1", isLocal: true),
ReplicaRow(role: "SECONDARY", replica: "NODE2"),
}, Ct);

Assert.Empty(h.Deliverer.Outcomes);

await e.ApplyAgReplicaHealthAsync(200002, "PRI", new[]
{
ReplicaRow(role: "SECONDARY", replica: "NODE1", isLocal: true),
ReplicaRow(role: "SECONDARY", replica: "NODE2"),
}, Ct);

var fired = Assert.Single(h.Deliverer.Outcomes);
Assert.Equal("AG Failover", fired.MetricName);
}

[Fact]
public async Task AgReplicaDisconnected_RefireOff_IsAPureEdge()
{
var h = new Harness();
var e = h.Build();

await e.ApplyAgReplicaHealthAsync(ServerId, Name, new[] { ReplicaRow(connected: "CONNECTED") }, Ct);
await e.ApplyAgReplicaHealthAsync(ServerId, Name, new[] { ReplicaRow(connected: "DISCONNECTED") }, Ct);
Assert.Single(h.Deliverer.Outcomes);

/* Default is off, so a replica down for a week still announces exactly once. */
h.Now = h.Now.AddHours(8);
await e.ApplyAgReplicaHealthAsync(ServerId, Name, new[] { ReplicaRow(connected: "DISCONNECTED") }, Ct);
Assert.Single(h.Deliverer.Outcomes);
}

[Fact]
public async Task AgReplicaDisconnected_Refire_ReAnnouncesUnderTheSameMetricName_ThenReconnectClearsTheClock()
{
var h = new Harness { AgDisconnectRefireMinutes = 10 };
var e = h.Build();

await e.ApplyAgReplicaHealthAsync(ServerId, Name, new[] { ReplicaRow(connected: "CONNECTED") }, Ct);
await e.ApplyAgReplicaHealthAsync(ServerId, Name, new[] { ReplicaRow(connected: "DISCONNECTED") }, Ct);
Assert.Single(h.Deliverer.Outcomes);

/* Inside the window: quiet. */
h.Now = h.Now.AddMinutes(5);
await e.ApplyAgReplicaHealthAsync(ServerId, Name, new[] { ReplicaRow(connected: "DISCONNECTED") }, Ct);
Assert.Single(h.Deliverer.Outcomes);

/* Past it: re-announce under the SAME metric name, so webhook automation keyed on it re-triggers. */
h.Now = h.Now.AddMinutes(6);
await e.ApplyAgReplicaHealthAsync(ServerId, Name, new[] { ReplicaRow(connected: "DISCONNECTED") }, Ct);
Assert.Equal(2, h.Deliverer.Outcomes.Count);
Assert.Equal("AG Replica Disconnected", h.Deliverer.Outcomes[1].MetricName);
Assert.Contains("STILL disconnected", h.Deliverer.Outcomes[1].ShortMessage, StringComparison.Ordinal);

/* Reconnect clears the clock, so a later outage starts a fresh episode rather than re-firing late. */
h.Now = h.Now.AddMinutes(1);
await e.ApplyAgReplicaHealthAsync(ServerId, Name, new[] { ReplicaRow(connected: "CONNECTED") }, Ct);
Assert.Equal(3, h.Deliverer.Outcomes.Count);
Assert.Equal("AG Replica Reconnected", h.Deliverer.Outcomes[2].MetricName);
}

/* ---------------- #991: gating and forget ---------------- */

[Fact]
Expand Down Expand Up @@ -1560,30 +1678,33 @@ public async Task AgAlerts_ThresholdsAreReadLive_SoAStoreEditTakesEffectOnTheNex
}

[Fact]
public async Task AgAlerts_Forget_DropsCompositeKeyedState_SoAReAddedServerReBaselines()
public async Task AgAlerts_Forget_ReleasesTheAgClaim_ButKeepsTheGroupsEdgeState()
{
var h = new Harness();
var e = h.Build();

/* Establish a role baseline and a standing sync-behind alert. */
await e.ApplyAgReplicaHealthAsync(ServerId, Name, new[] { ReplicaRow(role: "PRIMARY") }, Ct);
await e.ApplyAgDatabaseHealthAsync(ServerId, Name, new[] { DatabaseRow(lagSeconds: 600) }, Ct);
Assert.Single(h.Deliverer.Outcomes);
const int nodeA = 100001;
const int nodeB = 100002;

/* Removed from the monitored set. AG state is keyed by a COMPOSITE, so the per-server TryRemove that
clears the other conditions cannot reach it — Forget has to sweep the prefix. */
e.Forget(ServerId);
/* NODE-A claims the AG and establishes a role baseline; NODE-B sees the same AG but defers. */
await e.ApplyAgReplicaHealthAsync(nodeA, "NODE-A", new[] { ReplicaRow(role: "PRIMARY") }, Ct);
await e.ApplyAgReplicaHealthAsync(nodeB, "NODE-B", new[] { ReplicaRow(role: "PRIMARY") }, Ct);
Assert.Empty(h.Deliverer.Outcomes);

/* Re-added: the role is a fresh baseline (no phantom failover) ... */
await e.ApplyAgReplicaHealthAsync(ServerId, Name, new[] { ReplicaRow(role: "SECONDARY") }, Ct);
Assert.Single(h.Deliverer.Outcomes);
/* NODE-A is removed from monitoring. Its CLAIM is released so a survivor can take over, but the
AG's edge state is deliberately kept: the group still exists and NODE-B is still watching it.
Dropping the state here would re-baseline a live group and swallow the next failover. */
e.Forget(nodeA);

/* ... and the sync-behind episode starts over rather than sitting inside the dropped cooldown stamp. */
await e.ApplyAgDatabaseHealthAsync(ServerId, Name, new[] { DatabaseRow(lagSeconds: 600) }, Ct);
Assert.Equal(2, h.Deliverer.Outcomes.Count);
/* NODE-B takes over and sees the role change against the state NODE-A left behind — so the failover
is reported exactly once, by the new owner, with the correct previous role. */
await e.ApplyAgReplicaHealthAsync(nodeB, "NODE-B", new[] { ReplicaRow(role: "SECONDARY") }, Ct);

/* Nothing lingered to resolve, so no stray recovery row was written for the forgotten episode. */
Assert.Empty(h.History.Records);
var fired = Assert.Single(h.Deliverer.Outcomes);
Assert.Equal("AG Failover", fired.MetricName);
Assert.Equal("PRIMARY", fired.ThresholdValue);
Assert.Equal("SECONDARY", fired.CurrentValue);
Assert.Equal(nodeB.ToString(System.Globalization.CultureInfo.InvariantCulture), fired.ServerKey);
}

[Fact]
Expand Down
Loading
Loading