diff --git a/CHANGELOG.md b/CHANGELOG.md
index c026bc49e9..1dac0aa791 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -111,6 +111,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- **Lite's portable ZIP is self-contained, which HALVED it** ([#2501]) - `Publish Lite` is now `-r win-x64 --self-contained` in both `build.yml` and `nightly.yml`, so neither Lite artifact has a .NET prerequisite any more and the failure [#2489] documented stops existing: a tester who unzips onto a stock Windows Server no longer meets the .NET host's bare `You must install .NET to run this application` before a line of our code runs. **The size went the opposite way from what bundling a runtime suggests.** The old publish was RID-agnostic, so it copied every platform its packages ship - **537 MB of `runtimes\` on a 565 MB tree** (osx 130, linux-x64 116, linux-arm64 70, win-arm64 56, then win-x86, musl, loongarch64 and riscv64), of which only the **52 MB `win-x64`** folder could ever load on Windows. `DuckDB.NET.Bindings.Full` is most of it, SkiaSharp and SqlClient behind it. Dropping ~485 MB of unloadable native payload beats the cost of bundling .NET, WPF and ASP.NET Core by roughly two to one: measured on one commit and one SDK, **565 MB tree / 212.7 MB zipped becomes 277 MB / 114.2 MB**. It matters most for the **nightly** ZIP, which is the UAT download and is not offered as a `Setup.exe` at all. **A RID-specific publish needed two more files than the flag.** `Lite/packages.lock.json` had only a `net10.0-windows7.0` target, and a RID restore adds `net10.0-windows7.0/win-x64` to it - after which the `dotnet restore --locked-mode` that BOTH workflows run before the publish fails `NU1004: the project's runtime identifiers have changed`, because locked mode compares the PROJECT's RID set (empty) against the lock file's (win-x64). Reproduced locally; that is a red CI run on every PR, not the future `--no-restore` trap it was filed as. The fix is `win-x64` in `PerformanceMonitorLite.csproj`, so the project itself asks for that graph and one committed lock file satisfies the RID-less locked-mode restore and the RID publish alike; `RuntimeIdentifiers` (plural) sets no RID on the build, so a plain `dotnet build` stays RID-agnostic and `Lite.Tests` is untouched. **SignPath needed nothing** - the `Lite` artifact-configuration slug already receives both shapes today, and the signed re-zip reads `signed/Lite/*`, inheriting whatever shape `publish/Lite` has. Auto-update is unaffected; the ZIP is not a Velopack channel. `LiteRuntimePrerequisiteDocsTests` went red on the flag alone (3 of its 7 facts) and was rewritten to state every claim BOTH ways round: [#2499]'s version asserted only that the docs DID name the runtimes, so two of its facts stayed green while the prose went stale. It now also derives the lock file's RID coverage from the `-r` flags in the workflows, and every new assertion was proven red with its fix reverted.
### Fixed
+- **Every read in the alert evaluation pass now carries an explicit deadline** ([#2874]) - all forty-five commands across the six alert-pass types set no `CommandTimeout`, so every one inherited Npgsql's undocumented 30s default. On 2026-09-04 the forced-plan read failed five times on the production store, each surfacing as "Exception while reading from stream" - which is how Npgsql renders its OWN deadline, the misdiagnosis [#2826] exists to prevent. Unlike [#2810] and [#2871], which sit under `DarlingWorker.s_analysisTimeout`'s 120s `CancelAfter`, **this pass has no enclosing budget at all**: `EvaluateAlertsAsync` is called with the plain stopping token, so the per-command deadline IS the pass budget multiplied by the number of sequential reads - 45 x 30s of worst-case exposure while the body holds one of only four fleet sweep permits and that server's relaunch is skipped throughout. The value is therefore the first in this family set BELOW what it inherited rather than above: 10s, bounded under by the measured worst case (the shipped queries timed on the production store's three busiest servers put the whole pass's cost in one read - the forced-plan check at 1,744.9ms cold over ~6.0GB, every other read under 3ms) and bounded over by the 30s `s_alertSweepInterval` the pass runs on, so one stalled read still lets it finish inside the interval that restarts it. The asymmetry is what justifies erring short: an exceeded deadline skips one alert check and logs it, retrying 30s later, while a long-running read starves fleet-wide collection. Every observed failure was killed AT the 30s ceiling, so the record is right-censored - 10s is chosen from the measured cost and the cadence, not fitted to a failure distribution the data cannot show. Pinned structurally over both construction shapes: [#2874]'s census counted only `new NpgsqlCommand(`, missing `NpgsqlDataSource.CreateCommand(sql)`, which inherits the same default and accounts for four of this group's own sites - the repo-wide recount is recorded on the issue.
- **A held Ctrl+V during a busy clipboard no longer spawns several 'Pasted Plan' tabs from one keypress** ([#2870]) - [#2837] moved the clipboard can't-open retry off a synchronous `Thread.Sleep` onto an awaited `Task.Delay`, which keeps the UI pump responsive but also yields the thread for the ~175 ms retry window. The old `Thread.Sleep` had incidentally serialized input, so a HELD Ctrl+V (OS key-repeat) could not dispatch a second paste until the first finished; with the async retry, on a transiently busy clipboard several paste KeyDown handlers can sit in their retry loops at once, and when the clipboard frees each succeeds and loads its own tab. Each paste surface now carries a `_pasteInProgress` re-entrancy flag: set synchronously before the awaited read and cleared in a `finally` that runs only after the load finishes, so it spans the clipboard read AND the off-thread plan parse (the loaders return `Task` and the paste handlers `await` them, not fire-and-forget) and a repeat paste arriving during either window is dropped (its key stays claimed via `e.Handled`). The same flag guards both the Ctrl+V handler and the Paste XML button on each surface, across all three front ends (Lite, the Darling viewer, and the deprecated Dashboard). A rare, benign burst that the [#2837] async change had newly made possible; source-pinned so the guard can't be dropped without failing the build.
- **Retention held by the rollup-coverage gate is now visible, instead of reading as a healthy job** ([#2813]) - the #1680/#1877 gate PAUSES a retention policy so it cannot drop history a rollup has never materialized, which is correct and is not being weakened here. But a paused job is not a failing one: it reports `total_failures = 0`, `last_run_status = Success` and a plausible last-run duration, and `StoreSelfMetrics` never recorded `j.scheduled`, so every stored metric read clean. On the production store five `query_store_stats` policies sat held for **16 days** while the raw tier grew to 19 chunks and 65 GB - 4.5x its own 4-day horizon, and the store's largest object - and the only signal was one WARNING per service start ([#2809]). `get_store_metrics` now reports every held policy LIVE from the catalog, with the tier's actual span and how many times its configured `drop_after` it is really holding, and a `Retention Held` self-alert fires on the same hourly sweep as its [#2136] sibling. Judged on the CONJUNCTION of held AND past-horizon, never on the paused flag alone: `EnsureRetentionPoliciesAsync` deliberately creates every policy paused, so flag-only would fire on every fresh store at every start, and retention drops whole CHUNKS, so a 4-day policy with 1-day chunks legitimately sits at 1.25x while working perfectly. The ratio is deliberately not a hold duration - nothing records when a policy was paused - and measuring the consequence instead is what makes one threshold serve a 4-day raw tier and a 35-day baseline tier alike. Warning at 2.0x (clear of the 1.25x chunk floor, and ~8 days on the raw tier rather than the 16 it actually took), Critical at 4.0x, so the motivating incident reads Critical. The check never ACTS: arming a held policy drops the only copy of the history it holds, which is exactly what the gate prevents, so the release stays the `--backfill-rollups` operator decision and this makes the need for one visible rather than silent. Verified against a genuinely held state on TimescaleDB 2.28.1 - 19 chunks under a 4-day policy, measured 4.52x - by running the shipped reader itself, with the span normalization proven session-independent under UTC, UTC+14 and UTC-7. Persisting `scheduled` into the metrics series needs a migration rung and is tracked separately.
- **Every command in the analysis pass now carries an explicit deadline, not just the fact collector's** ([#2871]) - [#2810] gave `PgFactCollector`'s thirty-one commands a deliberate `CommandTimeout` and left the other THIRTY-TWO in the same 120s pass inheriting Npgsql's undocumented 30s default, and one of them was failing in production: on the dogfood box the `io_latency` baseline timed out nineteen times over two days and every sample landed between **30.1s and 31.4s**. A fixed wall, not a variable-duration fault and not the 15s connection timeout - the shipped query (dumped from the built assembly, not retyped) runs in **~1.6s** against the live store on the three busiest servers, so it crosses the ceiling only when the store stalls, which is why it read as intermittent. [#2820] had already made that query 5.6x faster (23.7s to 4.2s) and [#2826] had made the failure visible; neither touched what it was failing against, exactly as [#2827] and [#2826] had not touched [#2810]'s. The cost when it fires is larger than one pass: `GetOrComputeBaselinesAsync` assigns `_cache[cacheKey]` unconditionally, so a null result is CACHED and a single stall silences that (server, metric) pair's anomaly detection for the full 1-hour `CacheTtl`. **60s, bounded on both sides.** Below: the failure record is RIGHT-CENSORED - every run was killed at the ceiling, so nothing says whether it wanted 35s or 300s, and the value is therefore chosen as twice what it had rather than as a fitted number the data cannot support. Above - the bound that actually binds: `DarlingWorker.s_analysisTimeout` gives the pass 120s and the fact collector's thirty-one reads, up to eleven baseline computations per server, the anomaly detector and the drill-down all share it, so a per-command deadline near that figure would let ONE stalled command consume the pass and cost the server every other fact - strictly worse than the failure being fixed, which loses one metric. It matches `PgFactCollector.FactCommandTimeoutSeconds` deliberately, because two numbers in one budget would have to be reasoned about together every time either moved. Swept by SHAPE across the whole assembly rather than the one reported site (the [#2344] and `PgStatementText` mistake, twice paid for): `PgAnomalyDetector` 12, `PgFindingStore` 8, `PgDrillDownCollector` 10, `PgBaselineProvider` 1, `DarlingAnalysisService` 1. The drill-down's ten take that collector's OWN existing `DrillDownCommandTimeoutSeconds = 30` rather than a new value, so an author's deliberate choice on its other eight sites is not silently overridden. Pinned structurally over the pass - every command constructed in the assembly sets a deadline, and the value stays inside the justified band - and proven red three ways, each failing a DIFFERENT assertion: reverting the reported site, restoring 30s, and raising to the full 120s budget. Repo-wide this leaves 133 untimed `NpgsqlCommand` sites outside the analysis pass, which have different budgets and are filed separately rather than swept in here.
@@ -3206,4 +3207,5 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
[#2837]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2837
[#2870]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2870
[#2864]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2864
+[#2874]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2874
[#2876]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2876
diff --git a/Darling/Darling.Tests/AlertPassCommandTimeoutTests.cs b/Darling/Darling.Tests/AlertPassCommandTimeoutTests.cs
new file mode 100644
index 0000000000..40c336772d
--- /dev/null
+++ b/Darling/Darling.Tests/AlertPassCommandTimeoutTests.cs
@@ -0,0 +1,329 @@
+/*
+ * Copyright (c) 2026 Erik Darling, Darling Data LLC
+ *
+ * This file is part of the SQL Server Performance Monitor.
+ *
+ * Licensed under the MIT License. See LICENSE file in the project root for full license information.
+ */
+
+using System.Collections.Generic;
+using System.IO;
+using System.Linq;
+using System.Runtime.CompilerServices;
+using System.Text.RegularExpressions;
+using PerformanceMonitor.Darling.Service;
+using Xunit;
+
+namespace Darling.Tests;
+
+///
+/// Every read in the alert evaluation pass must carry an EXPLICIT command deadline (#2874).
+///
+/// All forty-five commands across the six alert-pass types ran with no
+/// CommandTimeout, so every one inherited Npgsql's undocumented 30 s default. On 2026-09-04
+/// the forced-plan read failed five times on the production store, each surfacing as "Exception
+/// while reading from stream" — how Npgsql renders its own deadline, and the exact misdiagnosis
+/// #2826 exists to prevent.
+///
+/// Why this pin matches TWO construction shapes. #2874's census counted 133 untimed
+/// sites by scanning for new NpgsqlCommand(. That shape is not the only one: an
+/// NpgsqlDataSource also hands out commands via CreateCommand(sql), which inherits the
+/// same default and which the census therefore missed entirely — four of this group's own sites are
+/// that shape, in DarlingPostgresAlertReadAdapter. A pin keyed on one spelling would have
+/// declared this family clean while a fifth of one file stayed on the inherited default, which is
+/// the #2786 failure exactly: a guard that names the arm it was written for. Both shapes are matched
+/// here, and the repo-wide recount is recorded on #2874.
+///
+/// The VALUE is pinned separately below, and is bounded on both sides for reasons that do not
+/// transfer from the two closed passes — see DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds.
+/// The short version: this pass has no enclosing CancelAfter, so the per-command deadline is
+/// the pass budget times the number of sequential reads, and it is deliberately set BELOW what it
+/// inherited rather than above.
+///
+public sealed class AlertPassCommandTimeoutTests
+{
+ ///
+ /// The six types that make up one alert evaluation pass. Named explicitly rather than globbed,
+ /// because "runs inside EvaluateAlertsAsync" is a budget boundary that no filename pattern
+ /// expresses — a future file in this directory may belong to a different budget. That is not
+ /// hypothetical: PgPlanForceActionStore looks like a member of this family and is not one.
+ /// Its only caller is PlanForceBot.RunAfterAnalysisAsync, dispatched as the analysis pass's
+ /// post-pass hook over the plain stopping token, so it shares the unbudgeted shape but runs on the
+ /// analysis interval rather than this pass's 30 s cadence — a different upper bound, and therefore
+ /// a different group.
+ ///
+ private static readonly string[] s_alertPassSources =
+ {
+ "DarlingAlertReadAdapter.cs",
+ "DarlingPostgresAlertReadAdapter.cs",
+ "PgAlertStateStore.cs",
+ "PgMuteRuleStore.cs",
+ "PgAlertHistoryStore.cs",
+ "DarlingSelfAlertEvaluator.cs",
+ };
+
+ ///
+ /// Both ways a command is constructed in this codebase. new NpgsqlCommand( is the shape
+ /// #2810's pin matched; .CreateCommand( is the one it did not, and which hid four sites
+ /// from #2874's census.
+ ///
+ private static readonly Regex s_commandCtor = new(
+ @"new NpgsqlCommand\s*\(|\.CreateCommand\s*\(",
+ RegexOptions.Compiled | RegexOptions.CultureInvariant);
+
+ private static readonly Regex s_setsTimeout = new(
+ @"CommandTimeout\s*=",
+ RegexOptions.Compiled | RegexOptions.CultureInvariant);
+
+ [Fact]
+ public void EveryAlertPassCommand_SetsAnExplicitDeadline()
+ {
+ var offenders = new List();
+ var total = 0;
+
+ foreach (var path in AlertPassSources())
+ {
+ var text = File.ReadAllText(path);
+
+ foreach (Match ctor in s_commandCtor.Matches(text))
+ {
+ total++;
+
+ /* Scan to the END OF THE STATEMENT, not a fixed number of lines.
+
+ A line window cannot work here and the first draft of this pin proved it: these
+ sites embed verbatim SQL, so the construction routinely spans twenty-plus lines and
+ the initializer that carries the deadline sits past any window small enough not to
+ run into the following member. A window wide enough to catch it would instead read
+ the NEXT command's deadline and call an untimed site clean — the failure that
+ actually matters, because it reports success on the defect.
+
+ The statement span is exact and needs no tuning: the deadline is either an object
+ initializer on the construction, or — for the CreateCommand shape, whose method
+ result cannot take one — the statement immediately after it, so two statements are
+ examined. */
+ var span = StatementSpanFrom(text, ctor.Index, statements: 2);
+
+ if (!s_setsTimeout.IsMatch(span))
+ {
+ var line = text.Take(ctor.Index).Count(c => c == '\n') + 1;
+ offenders.Add($"{Path.GetFileName(path)}:{line}");
+ }
+ }
+ }
+
+ Assert.True(total > 0, "the alert-pass scan matched no command constructions at all");
+
+ Assert.True(
+ offenders.Count == 0,
+ $"{offenders.Count} alert-pass command(s) inherit Npgsql's 30s default instead of setting "
+ + $"{nameof(DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds)}: "
+ + string.Join(", ", offenders));
+ }
+
+ ///
+ /// The value, bounded on both sides.
+ ///
+ /// ABOVE the measured worst case: the shipped queries timed against the production store's
+ /// three busiest servers put the whole pass's cost in one read — the forced-plan check at
+ /// 1,744.9 ms cold over ~6.0 GB, with every other read under 3 ms. A floor of 5 s keeps real
+ /// headroom over that.
+ ///
+ /// BELOW the 30 s s_alertSweepInterval this pass runs on, so one stalled read still
+ /// leaves the pass able to finish inside the interval that restarts it — and so the value stays
+ /// under the default it replaces, which is the whole point of this change.
+ ///
+ [Fact]
+ public void TheAlertPassDeadline_StaysInsideItsJustifiedBand()
+ {
+ var seconds = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds;
+
+ Assert.True(
+ seconds >= 5,
+ $"alert-pass deadline {seconds}s is at or under the measured 1.7s worst case with no "
+ + "meaningful headroom — a stall a little worse than normal would fail the read");
+
+ Assert.True(
+ seconds < 30,
+ $"alert-pass deadline {seconds}s is at or above Npgsql's inherited 30s default, so it "
+ + "buys nothing: this pass has no enclosing CancelAfter, and the per-command value is "
+ + "the pass budget times the number of sequential reads");
+ }
+
+ ///
+ /// The scanner above is what forty-five sites' correctness is asserted through, so its own blind
+ /// spots are worth pinning: a false positive here fails a green build on correct code.
+ ///
+ /// Both cases are shapes that actually occur. The first is the CreateCommand site
+ /// with an interposed explanatory comment containing a semicolon — the exact gap where this
+ /// codebase's style puts one. The second is the verbatim-SQL construction, where the deadline sits
+ /// twenty-odd lines below the opening and past any fixed line window.
+ ///
+ [Theory]
+ [InlineData(
+ "var command = _postgres.CreateCommand(Sql);\n"
+ + "/* set separately here; a method result cannot take an initializer. */\n"
+ + "command.CommandTimeout = 10;\n",
+ true)]
+ [InlineData(
+ "var command = new NpgsqlCommand(@\"\nSELECT 1;\nSELECT 2;\n\", connection) { CommandTimeout = 10 };\n",
+ true)]
+ [InlineData(
+ "var command = _postgres.CreateCommand(Sql);\n"
+ + "await command.ExecuteNonQueryAsync();\n",
+ false)]
+ public void TheScanner_SeesADeadlineThroughCommentsAndVerbatimSql(string source, bool expectedTimed)
+ {
+ var ctor = s_commandCtor.Match(source);
+ Assert.True(ctor.Success, "the fixture did not contain a command construction");
+
+ var span = StatementSpanFrom(source, ctor.Index, statements: 2);
+
+ Assert.Equal(expectedTimed, s_setsTimeout.IsMatch(span));
+ }
+
+ ///
+ /// The text from through the end of the Nth following statement,
+ /// counting only semicolons that sit outside string literals, outside comments, and outside any
+ /// nesting. Verbatim SQL in this family contains both semicolons and quote characters, so a naive
+ /// scan for ';' stops in the middle of a query and reports every multi-line construction as
+ /// untimed.
+ ///
+ /// Comments are skipped for the same reason and it is not hypothetical: the two-statement
+ /// window exists for the CreateCommand shape, whose deadline is the statement AFTER the
+ /// construction, and this codebase's style actively encourages an explanatory comment in exactly
+ /// that gap. A semicolon inside one would end the span early and report a correctly-timed site as
+ /// an offender — a false positive in the guard, which is worse than a miss because it fails a
+ /// green build and trains people to distrust the pin.
+ ///
+ private static string StatementSpanFrom(string text, int start, int statements)
+ {
+ var depth = 0;
+ var seen = 0;
+ var i = start;
+
+ while (i < text.Length)
+ {
+ var c = text[i];
+
+ if (c == '@' && i + 1 < text.Length && text[i + 1] == '"')
+ {
+ i = SkipVerbatimString(text, i + 2);
+ continue;
+ }
+
+ if (c == '"')
+ {
+ i = SkipRegularString(text, i + 1);
+ continue;
+ }
+
+ if (c == '/' && i + 1 < text.Length && text[i + 1] == '/')
+ {
+ var nl = text.IndexOf('\n', i);
+ i = nl < 0 ? text.Length : nl + 1;
+ continue;
+ }
+
+ if (c == '/' && i + 1 < text.Length && text[i + 1] == '*')
+ {
+ var end = text.IndexOf("*/", i + 2, System.StringComparison.Ordinal);
+ i = end < 0 ? text.Length : end + 2;
+ continue;
+ }
+
+ if (c is '(' or '[' or '{')
+ {
+ depth++;
+ }
+ else if (c is ')' or ']' or '}')
+ {
+ depth--;
+ }
+ else if (c == ';' && depth <= 0 && ++seen >= statements)
+ {
+ return text[start..(i + 1)];
+ }
+
+ i++;
+ }
+
+ return text[start..];
+ }
+
+ private static int SkipVerbatimString(string text, int i)
+ {
+ while (i < text.Length)
+ {
+ if (text[i] == '"')
+ {
+ /* "" is an escaped quote inside a verbatim string, not the end of it. */
+ if (i + 1 < text.Length && text[i + 1] == '"')
+ {
+ i += 2;
+ continue;
+ }
+
+ return i + 1;
+ }
+
+ i++;
+ }
+
+ return i;
+ }
+
+ private static int SkipRegularString(string text, int i)
+ {
+ while (i < text.Length)
+ {
+ if (text[i] == '\\')
+ {
+ i += 2;
+ continue;
+ }
+
+ if (text[i] == '"')
+ {
+ return i + 1;
+ }
+
+ i++;
+ }
+
+ return i;
+ }
+
+ private static IEnumerable AlertPassSources()
+ {
+ var dir = Path.Combine(RepoRoot(), "Darling", "PerformanceMonitor.Darling.Service");
+
+ var paths = s_alertPassSources
+ .Select(f => Path.Combine(dir, f))
+ .ToArray();
+
+ /* A renamed or moved alert-pass type must fail loudly here rather than silently shrinking the
+ scan to the files that still resolve — an empty or partial sweep is how a guard starts
+ reporting clean on code it no longer reads. */
+ foreach (var path in paths)
+ {
+ Assert.True(File.Exists(path), $"alert-pass source not found: {path}");
+ }
+
+ return paths;
+ }
+
+ private static string RepoRoot([CallerFilePath] string thisFile = "")
+ {
+ var dir = Path.GetDirectoryName(thisFile)!;
+ while (dir is not null
+ && !File.Exists(Path.Combine(dir, "PerformanceMonitor.sln"))
+ && !Directory.Exists(Path.Combine(dir, ".git")))
+ {
+ dir = Path.GetDirectoryName(dir);
+ }
+
+ Assert.NotNull(dir);
+ return dir!;
+ }
+}
diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingAlertReadAdapter.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingAlertReadAdapter.cs
index cf4208365a..686bd0a6bc 100644
--- a/Darling/PerformanceMonitor.Darling.Service/DarlingAlertReadAdapter.cs
+++ b/Darling/PerformanceMonitor.Darling.Service/DarlingAlertReadAdapter.cs
@@ -38,6 +38,56 @@ namespace PerformanceMonitor.Darling.Service;
///
public sealed class DarlingAlertReadAdapter : IAlertReadAdapter
{
+ ///
+ /// The explicit command deadline for EVERY read in the alert evaluation pass (#2874).
+ ///
+ /// All forty-five commands across the six alert-pass types ran with no
+ /// CommandTimeout, so every one inherited Npgsql's undocumented 30 s default. Nobody chose
+ /// 30 s; it was simply what happened. On 2026-09-04 the forced-plan read failed five times on the
+ /// production store, each time surfacing as "Exception while reading from stream" — which is how
+ /// Npgsql renders its OWN deadline, and which read literally says the network broke (the same
+ /// misdiagnosis #2826 exists to prevent).
+ ///
+ /// Why this pass needed its own number rather than the 60 s #2810 and #2871 chose.
+ /// Those two sit under DarlingWorker.s_analysisTimeout, a 120 s CancelAfter that
+ /// bounds the whole pass however long an individual command runs. This pass has no enclosing
+ /// budget at all — EvaluateAlertsAsync is called with the plain stopping token — so the
+ /// per-command deadline IS the pass budget, multiplied by however many reads run in sequence.
+ /// At the inherited 30 s that is 45 x 30 s of worst-case exposure while the body holds one of only
+ /// fleet permits, and the sweep skips
+ /// relaunch for that server the whole time. Copying 60 s here would have doubled it.
+ ///
+ /// Bounded below by measurement: the shipped queries were timed against the
+ /// production store on its three busiest servers, cold and warm. The whole pass is dominated by
+ /// one read — the forced-plan check at 1,744.9 ms cold, scanning ~6.0 GB of
+ /// query_store_stats; every other read in the family lands under 3 ms. Ten seconds is 5.7x
+ /// that worst case, so it absorbs a substantial stall rather than only the happy path.
+ ///
+ /// Bounded above by the cadence this pass runs on: s_alertSweepInterval is
+ /// 30 s, so one stalled read must still leave the pass able to finish inside the interval that
+ /// will start it again. Ten seconds keeps a single stall well inside that, and caps the unbudgeted
+ /// worst case at 45 x 10 s instead of 45 x 30 s.
+ ///
+ /// The asymmetry is why erring SHORT is right here, and it is the reverse of #2810.
+ /// A read that exceeds this deadline skips one alert check and logs it; the next pass runs 30 s
+ /// later, so the cost is one cycle of delay on one alert. A read that runs long holds a fleet
+ /// sweep permit and delays collection for every other server queued behind it. The recoverable
+ /// failure is strictly cheaper than the unrecoverable one, so this is the first value in the
+ /// family set BELOW what it inherited rather than above it.
+ ///
+ /// What the data cannot say. Every observed failure was killed AT the 30 s ceiling,
+ /// so the record is right-censored: nothing here establishes whether a stalled read wanted 35 s or
+ /// 300 s. Ten seconds is chosen from the measured cost and the cadence above, NOT fitted to the
+ /// failure distribution — a number claiming to fit that data would be invented.
+ ///
+ /// What this does NOT cover. PgPlanForceActionStore sits beside these
+ /// types and is not one of them: its only caller is PlanForceBot.RunAfterAnalysisAsync,
+ /// dispatched as the analysis pass's post-pass hook over the plain stopping token. It shares
+ /// the unbudgeted shape but runs on the analysis interval, not this pass's 30 s cadence, so the
+ /// upper bound derived above does not apply to it and it is left for its own group (#2874).
+ ///
+ internal const int AlertPassCommandTimeoutSeconds = 10;
+
private readonly NpgsqlDataSource _postgres;
private readonly Func? _runningJobsCadenceMinutes;
private readonly Func? _blockingSnapshotCadenceMinutes;
@@ -121,7 +171,7 @@ public async Task> GetRecentBlockedProcessReportsAs
await using (var connection = await _postgres.OpenConnectionAsync(cancellationToken))
{
- using (var command = new NpgsqlCommand(BlockedProcessReportsSql, connection))
+ using (var command = new NpgsqlCommand(BlockedProcessReportsSql, connection) { CommandTimeout = AlertPassCommandTimeoutSeconds })
{
command.Parameters.AddWithValue(serverId);
command.Parameters.AddWithValue(startTime);
@@ -146,7 +196,7 @@ public async Task> GetRecentBlockedProcessReportsAs
}
}
- using (var command = new NpgsqlCommand(DmvBlockingSnapshotsSql, connection))
+ using (var command = new NpgsqlCommand(DmvBlockingSnapshotsSql, connection) { CommandTimeout = AlertPassCommandTimeoutSeconds })
{
command.Parameters.AddWithValue(serverId);
command.Parameters.AddWithValue(startTime);
@@ -213,7 +263,7 @@ FROM dmv_blocking_snapshots
var serverId = ParseServerKey(serverKey);
await using var connection = await _postgres.OpenConnectionAsync(cancellationToken);
- using var command = new NpgsqlCommand(CurrentBlockingWaitSql, connection);
+ using var command = new NpgsqlCommand(CurrentBlockingWaitSql, connection) { CommandTimeout = AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(serverId);
using var reader = await command.ExecuteReaderAsync(cancellationToken);
@@ -259,7 +309,7 @@ public async Task> GetRecentDeadlocksAsync(
var items = new List();
await using var connection = await _postgres.OpenConnectionAsync(cancellationToken);
- using var command = new NpgsqlCommand(DeadlocksSql, connection);
+ using var command = new NpgsqlCommand(DeadlocksSql, connection) { CommandTimeout = AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(serverId);
command.Parameters.AddWithValue(startTime);
command.Parameters.AddWithValue(endTime);
@@ -309,7 +359,7 @@ public async Task> GetPoisonWaitDeltasAsync(
var items = new List();
await using var connection = await _postgres.OpenConnectionAsync(cancellationToken);
- using var command = new NpgsqlCommand(PoisonWaitsSql, connection);
+ using var command = new NpgsqlCommand(PoisonWaitsSql, connection) { CommandTimeout = AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(serverId);
command.Parameters.AddWithValue(NaiveUtcNow().AddMinutes(-10));
@@ -406,7 +456,7 @@ public async Task> GetLongRunningQueriesAsync(
var items = new List();
await using var connection = await _postgres.OpenConnectionAsync(cancellationToken);
- using var command = new NpgsqlCommand(sql, connection);
+ using var command = new NpgsqlCommand(sql, connection) { CommandTimeout = AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(serverId);
command.Parameters.AddWithValue(thresholdMs);
command.Parameters.AddWithValue(maxResults);
@@ -512,7 +562,7 @@ public async Task> GetDatabaseFileGrowthAsync(
var items = new List();
await using var connection = await _postgres.OpenConnectionAsync(cancellationToken);
- using var command = new NpgsqlCommand(DatabaseFileGrowthSql, connection);
+ using var command = new NpgsqlCommand(DatabaseFileGrowthSql, connection) { CommandTimeout = AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(serverId);
command.Parameters.AddWithValue(windowStart);
@@ -568,7 +618,7 @@ public async Task> GetVolumeFreeSpaceAsync(
var items = new List();
await using var connection = await _postgres.OpenConnectionAsync(cancellationToken);
- using var command = new NpgsqlCommand(VolumeFreeSpaceSql, connection);
+ using var command = new NpgsqlCommand(VolumeFreeSpaceSql, connection) { CommandTimeout = AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(serverId);
using var reader = await command.ExecuteReaderAsync(cancellationToken);
@@ -623,7 +673,7 @@ public async Task> GetPvsPressureAsync(
var items = new List();
await using var connection = await _postgres.OpenConnectionAsync(cancellationToken);
- using var command = new NpgsqlCommand(PvsPressureSql, connection);
+ using var command = new NpgsqlCommand(PvsPressureSql, connection) { CommandTimeout = AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(serverId);
using var reader = await command.ExecuteReaderAsync(cancellationToken);
@@ -669,7 +719,7 @@ ORDER BY collection_time DESC
var serverId = ParseServerKey(serverKey);
await using var connection = await _postgres.OpenConnectionAsync(cancellationToken);
- using var command = new NpgsqlCommand(TempDbSpaceSql, connection);
+ using var command = new NpgsqlCommand(TempDbSpaceSql, connection) { CommandTimeout = AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(serverId);
using var reader = await command.ExecuteReaderAsync(cancellationToken);
@@ -731,7 +781,7 @@ public async Task GetAnomalousJobsAsync(
per-run cooldown key expires each pass, so a stale snapshot re-fires the same historical run
every cooldown, forever. Same rule as Lite's adapter; parity is the point. */
using (var snapshotProbe = new NpgsqlCommand(
- "SELECT MAX(collection_time) FROM running_jobs WHERE server_id = $1", connection))
+ "SELECT MAX(collection_time) FROM running_jobs WHERE server_id = $1", connection) { CommandTimeout = AlertPassCommandTimeoutSeconds })
{
snapshotProbe.Parameters.AddWithValue(serverId);
var snapshot = await snapshotProbe.ExecuteScalarAsync(cancellationToken);
@@ -743,7 +793,7 @@ public async Task GetAnomalousJobsAsync(
}
}
- using var command = new NpgsqlCommand(AnomalousJobsSql, connection);
+ using var command = new NpgsqlCommand(AnomalousJobsSql, connection) { CommandTimeout = AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(serverId);
command.Parameters.AddWithValue(thresholdPercent);
@@ -946,7 +996,7 @@ public async Task> GetDatabaseStatesAsync(
var items = new List();
await using var connection = await _postgres.OpenConnectionAsync(cancellationToken);
- using (var seed = new NpgsqlCommand(SeedDatabaseStateExpectedSql, connection))
+ using (var seed = new NpgsqlCommand(SeedDatabaseStateExpectedSql, connection) { CommandTimeout = AlertPassCommandTimeoutSeconds })
{
seed.Parameters.AddWithValue(serverId);
await seed.ExecuteNonQueryAsync(cancellationToken);
@@ -956,13 +1006,13 @@ public async Task> GetDatabaseStatesAsync(
for a database that has none, this un-learns one the database has since outgrown. Both run before
the read, so a poisoned expectation is corrected on the cycle that notices it rather than firing
once more first. */
- using (var heal = new NpgsqlCommand(HealDatabaseStateBaselineToOnlineSql, connection))
+ using (var heal = new NpgsqlCommand(HealDatabaseStateBaselineToOnlineSql, connection) { CommandTimeout = AlertPassCommandTimeoutSeconds })
{
heal.Parameters.AddWithValue(serverId);
await heal.ExecuteNonQueryAsync(cancellationToken);
}
- using (var prune = new NpgsqlCommand(PruneDatabaseStateExpectedSql, connection))
+ using (var prune = new NpgsqlCommand(PruneDatabaseStateExpectedSql, connection) { CommandTimeout = AlertPassCommandTimeoutSeconds })
{
prune.Parameters.AddWithValue(serverId);
await prune.ExecuteNonQueryAsync(cancellationToken);
@@ -972,13 +1022,13 @@ once more first. */
one carried over from a restart (#2166). A database cleared here is one that is back at its
expected state, so it cannot appear in the deviation read below either way — the ordering matters
for the NEXT deviation, not this one. */
- using (var clearRecovered = new NpgsqlCommand(ClearRecoveredDatabaseStateAlertsSql, connection))
+ using (var clearRecovered = new NpgsqlCommand(ClearRecoveredDatabaseStateAlertsSql, connection) { CommandTimeout = AlertPassCommandTimeoutSeconds })
{
clearRecovered.Parameters.AddWithValue(serverId);
await clearRecovered.ExecuteNonQueryAsync(cancellationToken);
}
- using (var command = new NpgsqlCommand(DatabaseStateDeviationsSql, connection))
+ using (var command = new NpgsqlCommand(DatabaseStateDeviationsSql, connection) { CommandTimeout = AlertPassCommandTimeoutSeconds })
{
command.Parameters.AddWithValue(serverId);
using var reader = await command.ExecuteReaderAsync(cancellationToken);
@@ -1058,7 +1108,7 @@ public async Task> GetForcePlanFailuresAsync(
var items = new List();
await using var connection = await _postgres.OpenConnectionAsync(cancellationToken);
- using var command = new NpgsqlCommand(ForcePlanFailuresSql, connection);
+ using var command = new NpgsqlCommand(ForcePlanFailuresSql, connection) { CommandTimeout = AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(serverId);
using var reader = await command.ExecuteReaderAsync(cancellationToken);
diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingPostgresAlertReadAdapter.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingPostgresAlertReadAdapter.cs
index be2de0aefa..23ed7d0b7f 100644
--- a/Darling/PerformanceMonitor.Darling.Service/DarlingPostgresAlertReadAdapter.cs
+++ b/Darling/PerformanceMonitor.Darling.Service/DarlingPostgresAlertReadAdapter.cs
@@ -181,6 +181,7 @@ public async Task> GetPoisonWaitPressureAsync(
{
var rows = new List();
await using var command = _postgres.CreateCommand(PoisonWaitSql);
+ command.CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds;
command.Parameters.AddWithValue(serverId);
/* The evaluator's window, not this adapter's 2-hour Freshness: the window IS the denominator the
threshold normalizes against, so read and evaluation must agree on it or the "average backends
@@ -206,6 +207,7 @@ public async Task> GetWraparoundRiskAsync(
{
var rows = new List();
await using var command = _postgres.CreateCommand(WraparoundSql);
+ command.CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds;
command.Parameters.AddWithValue(serverId);
command.Parameters.AddWithValue(NaiveUtcNow() - Freshness);
await using var reader = await command.ExecuteReaderAsync(cancellationToken);
@@ -236,6 +238,7 @@ answer consistently off one query. 0 reads as "no window data" -> FreezingIsKeep
int serverId, CancellationToken cancellationToken = default)
{
await using var command = _postgres.CreateCommand(XminSql);
+ command.CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds;
command.Parameters.AddWithValue(serverId);
command.Parameters.AddWithValue(NaiveUtcNow() - Freshness);
await using var reader = await command.ExecuteReaderAsync(cancellationToken);
@@ -258,6 +261,7 @@ public async Task> GetReplicationSlotRiskAsync(
{
var rows = new List();
await using var command = _postgres.CreateCommand(SlotSql);
+ command.CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds;
command.Parameters.AddWithValue(serverId);
command.Parameters.AddWithValue(NaiveUtcNow() - Freshness);
await using var reader = await command.ExecuteReaderAsync(cancellationToken);
diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingSelfAlertEvaluator.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingSelfAlertEvaluator.cs
index bbad1a2e57..f5a642af06 100644
--- a/Darling/PerformanceMonitor.Darling.Service/DarlingSelfAlertEvaluator.cs
+++ b/Darling/PerformanceMonitor.Darling.Service/DarlingSelfAlertEvaluator.cs
@@ -2023,7 +2023,7 @@ AND status IN ('SUCCESS', 'SKIPPED')) A
FROM (SELECT log_id FROM collection_log WHERE server_id = $1 ORDER BY log_id DESC LIMIT $2) r) AS recent_runs,
(SELECT COUNT(*)
FROM (SELECT status FROM collection_log WHERE server_id = $1 ORDER BY log_id DESC LIMIT $2) r
- WHERE r.status IN ('SUCCESS', 'SKIPPED')) AS recent_success", connection);
+ WHERE r.status IN ('SUCCESS', 'SKIPPED')) AS recent_success", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(serverId);
command.Parameters.AddWithValue(recentWindow);
@@ -2067,7 +2067,7 @@ AND cl.collector_name IN ('deadlocks', 'blocked_process_report')
) AS x
WHERE x.n = 1
AND x.status = 'SESSION_MISSING'
-ORDER BY x.collector_name", connection);
+ORDER BY x.collector_name", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(serverId);
await using var reader = await command.ExecuteReaderAsync(cancellationToken);
@@ -2104,7 +2104,7 @@ SELECT 1
FROM agent_status
WHERE server_id = $1
AND agent_running
-)", connection);
+)", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(serverId);
return (bool)(await command.ExecuteScalarAsync(cancellationToken))!;
}
@@ -2125,7 +2125,7 @@ AND agent_running
FROM agent_status
WHERE server_id = $1
ORDER BY collection_time DESC
-LIMIT 1", connection);
+LIMIT 1", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(serverId);
await using var reader = await command.ExecuteReaderAsync(cancellationToken);
@@ -2172,7 +2172,7 @@ ORDER BY collection_time DESC
FROM ag_replica_states
WHERE server_id = $1
AND collection_time = (SELECT MAX(collection_time) FROM ag_replica_states WHERE server_id = $1)
-ORDER BY ag_name, replica_server_name", connection);
+ORDER BY ag_name, replica_server_name", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(serverId);
await using var reader = await command.ExecuteReaderAsync(cancellationToken);
@@ -2216,7 +2216,7 @@ FROM ag_replica_states
FROM ag_database_replica_states
WHERE server_id = $1
AND collection_time = (SELECT MAX(collection_time) FROM ag_database_replica_states WHERE server_id = $1)
-ORDER BY ag_name, database_name, replica_server_name", connection);
+ORDER BY ag_name, database_name, replica_server_name", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(serverId);
await using var reader = await command.ExecuteReaderAsync(cancellationToken);
diff --git a/Darling/PerformanceMonitor.Darling.Service/PgAlertHistoryStore.cs b/Darling/PerformanceMonitor.Darling.Service/PgAlertHistoryStore.cs
index 5c9b2ffdf9..505cf78611 100644
--- a/Darling/PerformanceMonitor.Darling.Service/PgAlertHistoryStore.cs
+++ b/Darling/PerformanceMonitor.Darling.Service/PgAlertHistoryStore.cs
@@ -62,7 +62,7 @@ duplication was the drift risk #1881 turned on. */
await using var connection = await _postgres.OpenConnectionAsync();
using var command = new NpgsqlCommand(@"
INSERT INTO config_alert_log (alert_time, server_id, server_name, metric_name, current_value, threshold_value, alert_sent, notification_type, send_error, muted, detail_text, context_json)
-VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11, $12)", connection);
+VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11, $12)", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(NaiveUtcNow());
command.Parameters.AddWithValue(serverId);
@@ -135,7 +135,7 @@ FROM config_alert_log
special characters escaped) by AlertContextSerializer.BuildDedupKeyLikePattern
rather than hand-concatenated here — see its doc comment for why. NULL context_json
rows fail the match either way. */
- + (dedupKey is null ? "" : "\nAND context_json LIKE $3 ESCAPE '\\'"), connection);
+ + (dedupKey is null ? "" : "\nAND context_json LIKE $3 ESCAPE '\\'"), connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(sid);
command.Parameters.AddWithValue(metricName);
if (dedupKey is not null)
diff --git a/Darling/PerformanceMonitor.Darling.Service/PgAlertStateStore.cs b/Darling/PerformanceMonitor.Darling.Service/PgAlertStateStore.cs
index a46c90b23d..6fe8e34056 100644
--- a/Darling/PerformanceMonitor.Darling.Service/PgAlertStateStore.cs
+++ b/Darling/PerformanceMonitor.Darling.Service/PgAlertStateStore.cs
@@ -53,7 +53,7 @@ SELECT watermark
FROM config_edge_trigger_watermarks
WHERE server_id = $1
AND metric_name = $2
-AND watermark_time IS NULL", connection);
+AND watermark_time IS NULL", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(ParseServerKey(serverKey));
command.Parameters.AddWithValue(metricName);
@@ -85,7 +85,7 @@ INSERT INTO config_edge_trigger_watermarks (server_id, metric_name, watermark, w
ON CONFLICT (server_id, metric_name) DO UPDATE SET
watermark = EXCLUDED.watermark,
watermark_time = NULL,
- updated_at = EXCLUDED.updated_at", connection);
+ updated_at = EXCLUDED.updated_at", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(ParseServerKey(serverKey));
command.Parameters.AddWithValue(metricName);
command.Parameters.AddWithValue(watermark);
@@ -109,7 +109,7 @@ SELECT watermark_time
FROM config_edge_trigger_watermarks
WHERE server_id = $1
AND metric_name = $2
-AND watermark_time IS NOT NULL", connection);
+AND watermark_time IS NOT NULL", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(ParseServerKey(serverKey));
command.Parameters.AddWithValue(FailedJobWatermarkMetric);
@@ -150,7 +150,7 @@ INSERT INTO config_edge_trigger_watermarks (server_id, metric_name, watermark, w
ON CONFLICT (server_id, metric_name) DO UPDATE SET
watermark = 0,
watermark_time = EXCLUDED.watermark_time,
- updated_at = EXCLUDED.updated_at", connection);
+ updated_at = EXCLUDED.updated_at", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(ParseServerKey(serverKey));
command.Parameters.AddWithValue(FailedJobWatermarkMetric);
command.Parameters.AddWithValue(DateTime.SpecifyKind(watermark, DateTimeKind.Unspecified));
@@ -190,7 +190,7 @@ UPDATE config.database_state_expected
SET last_alerted_state = $3,
last_alerted_at = (now() AT TIME ZONE 'UTC')
WHERE server_id = $1
-AND database_name = $2", connection);
+AND database_name = $2", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(ParseServerKey(serverKey));
command.Parameters.AddWithValue(databaseName);
command.Parameters.AddWithValue(effectiveState);
@@ -221,7 +221,7 @@ UPDATE config.database_state_expected
SET last_alerted_state = NULL,
last_alerted_at = NULL
WHERE server_id = $1
-AND database_name = $2", connection);
+AND database_name = $2", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(ParseServerKey(serverKey));
command.Parameters.AddWithValue(databaseName);
await command.ExecuteNonQueryAsync();
@@ -257,7 +257,7 @@ public async Task> LoadInci
SELECT dedup_key, total_occurrences, observed_window_count, incident_started_at, last_observed_at
FROM config.incident_occurrences
WHERE server_id = $1
-AND metric_name = $2", connection);
+AND metric_name = $2", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(ParseServerKey(serverKey));
command.Parameters.AddWithValue(metricName);
@@ -338,7 +338,7 @@ never admits a blank or null fingerprint (it passes those incidents through unke
DELETE FROM config.incident_occurrences
WHERE server_id = $1
AND metric_name = $2
-AND dedup_key <> ALL($3::text[])", connection, transaction))
+AND dedup_key <> ALL($3::text[])", connection, transaction) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds })
{
prune.Parameters.AddWithValue(serverId);
prune.Parameters.AddWithValue(metricName);
@@ -358,7 +358,7 @@ ON CONFLICT (server_id, metric_name, dedup_key) DO UPDATE SET
total_occurrences = EXCLUDED.total_occurrences,
observed_window_count = EXCLUDED.observed_window_count,
incident_started_at = EXCLUDED.incident_started_at,
- last_observed_at = EXCLUDED.last_observed_at", connection, transaction);
+ last_observed_at = EXCLUDED.last_observed_at", connection, transaction) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
upsert.Parameters.AddWithValue(serverId);
upsert.Parameters.AddWithValue(metricName);
upsert.Parameters.AddWithValue(dedupKeys);
diff --git a/Darling/PerformanceMonitor.Darling.Service/PgMuteRuleStore.cs b/Darling/PerformanceMonitor.Darling.Service/PgMuteRuleStore.cs
index 3f928d0234..e5c6dd285d 100644
--- a/Darling/PerformanceMonitor.Darling.Service/PgMuteRuleStore.cs
+++ b/Darling/PerformanceMonitor.Darling.Service/PgMuteRuleStore.cs
@@ -47,7 +47,7 @@ public async Task> LoadAllAsync()
server_name, metric_name, database_pattern,
query_text_pattern, wait_type_pattern, job_name_pattern
FROM config_mute_rules
-ORDER BY created_at_utc DESC", connection);
+ORDER BY created_at_utc DESC", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
using var reader = await command.ExecuteReaderAsync();
while (await reader.ReadAsync())
@@ -91,7 +91,7 @@ INSERT INTO config_mute_rules
(id, enabled, created_at_utc, expires_at_utc, reason,
server_name, metric_name, database_pattern,
query_text_pattern, wait_type_pattern, job_name_pattern)
-VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11)", connection);
+VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11)", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(rule.Id);
command.Parameters.AddWithValue(rule.Enabled);
command.Parameters.AddWithValue(Naive(rule.CreatedAtUtc));
@@ -119,7 +119,7 @@ UPDATE config_mute_rules SET
enabled = $2, expires_at_utc = $3, reason = $4,
server_name = $5, metric_name = $6, database_pattern = $7,
query_text_pattern = $8, wait_type_pattern = $9, job_name_pattern = $10
-WHERE id = $1", connection);
+WHERE id = $1", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(rule.Id);
command.Parameters.AddWithValue(rule.Enabled);
AddNullable(command, rule.ExpiresAtUtc);
@@ -136,7 +136,7 @@ UPDATE config_mute_rules SET
public async Task SetEnabledAsync(string ruleId, bool enabled)
{
await using var connection = await _postgres.OpenConnectionAsync();
- using var command = new NpgsqlCommand("UPDATE config_mute_rules SET enabled = $2 WHERE id = $1", connection);
+ using var command = new NpgsqlCommand("UPDATE config_mute_rules SET enabled = $2 WHERE id = $1", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(ruleId);
command.Parameters.AddWithValue(enabled);
await command.ExecuteNonQueryAsync();
@@ -145,7 +145,7 @@ public async Task SetEnabledAsync(string ruleId, bool enabled)
public async Task DeleteAsync(string ruleId)
{
await using var connection = await _postgres.OpenConnectionAsync();
- using var command = new NpgsqlCommand("DELETE FROM config_mute_rules WHERE id = $1", connection);
+ using var command = new NpgsqlCommand("DELETE FROM config_mute_rules WHERE id = $1", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(ruleId);
await command.ExecuteNonQueryAsync();
}
@@ -160,7 +160,7 @@ public async Task DeleteExpiredAsync(IReadOnlyList expiredIds)
await using var connection = await _postgres.OpenConnectionAsync();
foreach (var id in expiredIds)
{
- using var command = new NpgsqlCommand("DELETE FROM config_mute_rules WHERE id = $1", connection);
+ using var command = new NpgsqlCommand("DELETE FROM config_mute_rules WHERE id = $1", connection) { CommandTimeout = DarlingAlertReadAdapter.AlertPassCommandTimeoutSeconds };
command.Parameters.AddWithValue(id);
await command.ExecuteNonQueryAsync();
}