Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -111,6 +111,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- **Lite's portable ZIP is self-contained, which HALVED it** ([#2501]) - `Publish Lite` is now `-r win-x64 --self-contained` in both `build.yml` and `nightly.yml`, so neither Lite artifact has a .NET prerequisite any more and the failure [#2489] documented stops existing: a tester who unzips onto a stock Windows Server no longer meets the .NET host's bare `You must install .NET to run this application` before a line of our code runs. **The size went the opposite way from what bundling a runtime suggests.** The old publish was RID-agnostic, so it copied every platform its packages ship - **537 MB of `runtimes\` on a 565 MB tree** (osx 130, linux-x64 116, linux-arm64 70, win-arm64 56, then win-x86, musl, loongarch64 and riscv64), of which only the **52 MB `win-x64`** folder could ever load on Windows. `DuckDB.NET.Bindings.Full` is most of it, SkiaSharp and SqlClient behind it. Dropping ~485 MB of unloadable native payload beats the cost of bundling .NET, WPF and ASP.NET Core by roughly two to one: measured on one commit and one SDK, **565 MB tree / 212.7 MB zipped becomes 277 MB / 114.2 MB**. It matters most for the **nightly** ZIP, which is the UAT download and is not offered as a `Setup.exe` at all. **A RID-specific publish needed two more files than the flag.** `Lite/packages.lock.json` had only a `net10.0-windows7.0` target, and a RID restore adds `net10.0-windows7.0/win-x64` to it - after which the `dotnet restore --locked-mode` that BOTH workflows run before the publish fails `NU1004: the project's runtime identifiers have changed`, because locked mode compares the PROJECT's RID set (empty) against the lock file's (win-x64). Reproduced locally; that is a red CI run on every PR, not the future `--no-restore` trap it was filed as. The fix is `<RuntimeIdentifiers>win-x64</RuntimeIdentifiers>` in `PerformanceMonitorLite.csproj`, so the project itself asks for that graph and one committed lock file satisfies the RID-less locked-mode restore and the RID publish alike; `RuntimeIdentifiers` (plural) sets no RID on the build, so a plain `dotnet build` stays RID-agnostic and `Lite.Tests` is untouched. **SignPath needed nothing** - the `Lite` artifact-configuration slug already receives both shapes today, and the signed re-zip reads `signed/Lite/*`, inheriting whatever shape `publish/Lite` has. Auto-update is unaffected; the ZIP is not a Velopack channel. `LiteRuntimePrerequisiteDocsTests` went red on the flag alone (3 of its 7 facts) and was rewritten to state every claim BOTH ways round: [#2499]'s version asserted only that the docs DID name the runtimes, so two of its facts stayed green while the prose went stale. It now also derives the lock file's RID coverage from the `-r` flags in the workflows, and every new assertion was proven red with its fix reverted.

### Fixed
- **Every read in the alert evaluation pass now carries an explicit deadline** ([#2874]) - all forty-five commands across the six alert-pass types set no `CommandTimeout`, so every one inherited Npgsql's undocumented 30s default. On 2026-09-04 the forced-plan read failed five times on the production store, each surfacing as "Exception while reading from stream" - which is how Npgsql renders its OWN deadline, the misdiagnosis [#2826] exists to prevent. Unlike [#2810] and [#2871], which sit under `DarlingWorker.s_analysisTimeout`'s 120s `CancelAfter`, **this pass has no enclosing budget at all**: `EvaluateAlertsAsync` is called with the plain stopping token, so the per-command deadline IS the pass budget multiplied by the number of sequential reads - 45 x 30s of worst-case exposure while the body holds one of only four fleet sweep permits and that server's relaunch is skipped throughout. The value is therefore the first in this family set BELOW what it inherited rather than above: 10s, bounded under by the measured worst case (the shipped queries timed on the production store's three busiest servers put the whole pass's cost in one read - the forced-plan check at 1,744.9ms cold over ~6.0GB, every other read under 3ms) and bounded over by the 30s `s_alertSweepInterval` the pass runs on, so one stalled read still lets it finish inside the interval that restarts it. The asymmetry is what justifies erring short: an exceeded deadline skips one alert check and logs it, retrying 30s later, while a long-running read starves fleet-wide collection. Every observed failure was killed AT the 30s ceiling, so the record is right-censored - 10s is chosen from the measured cost and the cadence, not fitted to a failure distribution the data cannot show. Pinned structurally over both construction shapes: [#2874]'s census counted only `new NpgsqlCommand(`, missing `NpgsqlDataSource.CreateCommand(sql)`, which inherits the same default and accounts for four of this group's own sites - the repo-wide recount is recorded on the issue.
- **A held Ctrl+V during a busy clipboard no longer spawns several 'Pasted Plan' tabs from one keypress** ([#2870]) - [#2837] moved the clipboard can't-open retry off a synchronous `Thread.Sleep` onto an awaited `Task.Delay`, which keeps the UI pump responsive but also yields the thread for the ~175 ms retry window. The old `Thread.Sleep` had incidentally serialized input, so a HELD Ctrl+V (OS key-repeat) could not dispatch a second paste until the first finished; with the async retry, on a transiently busy clipboard several paste KeyDown handlers can sit in their retry loops at once, and when the clipboard frees each succeeds and loads its own tab. Each paste surface now carries a `_pasteInProgress` re-entrancy flag: set synchronously before the awaited read and cleared in a `finally` that runs only after the load finishes, so it spans the clipboard read AND the off-thread plan parse (the loaders return `Task` and the paste handlers `await` them, not fire-and-forget) and a repeat paste arriving during either window is dropped (its key stays claimed via `e.Handled`). The same flag guards both the Ctrl+V handler and the Paste XML button on each surface, across all three front ends (Lite, the Darling viewer, and the deprecated Dashboard). A rare, benign burst that the [#2837] async change had newly made possible; source-pinned so the guard can't be dropped without failing the build.
- **Retention held by the rollup-coverage gate is now visible, instead of reading as a healthy job** ([#2813]) - the #1680/#1877 gate PAUSES a retention policy so it cannot drop history a rollup has never materialized, which is correct and is not being weakened here. But a paused job is not a failing one: it reports `total_failures = 0`, `last_run_status = Success` and a plausible last-run duration, and `StoreSelfMetrics` never recorded `j.scheduled`, so every stored metric read clean. On the production store five `query_store_stats` policies sat held for **16 days** while the raw tier grew to 19 chunks and 65 GB - 4.5x its own 4-day horizon, and the store's largest object - and the only signal was one WARNING per service start ([#2809]). `get_store_metrics` now reports every held policy LIVE from the catalog, with the tier's actual span and how many times its configured `drop_after` it is really holding, and a `Retention Held` self-alert fires on the same hourly sweep as its [#2136] sibling. Judged on the CONJUNCTION of held AND past-horizon, never on the paused flag alone: `EnsureRetentionPoliciesAsync` deliberately creates every policy paused, so flag-only would fire on every fresh store at every start, and retention drops whole CHUNKS, so a 4-day policy with 1-day chunks legitimately sits at 1.25x while working perfectly. The ratio is deliberately not a hold duration - nothing records when a policy was paused - and measuring the consequence instead is what makes one threshold serve a 4-day raw tier and a 35-day baseline tier alike. Warning at 2.0x (clear of the 1.25x chunk floor, and ~8 days on the raw tier rather than the 16 it actually took), Critical at 4.0x, so the motivating incident reads Critical. The check never ACTS: arming a held policy drops the only copy of the history it holds, which is exactly what the gate prevents, so the release stays the `--backfill-rollups` operator decision and this makes the need for one visible rather than silent. Verified against a genuinely held state on TimescaleDB 2.28.1 - 19 chunks under a 4-day policy, measured 4.52x - by running the shipped reader itself, with the span normalization proven session-independent under UTC, UTC+14 and UTC-7. Persisting `scheduled` into the metrics series needs a migration rung and is tracked separately.
- **Every command in the analysis pass now carries an explicit deadline, not just the fact collector's** ([#2871]) - [#2810] gave `PgFactCollector`'s thirty-one commands a deliberate `CommandTimeout` and left the other THIRTY-TWO in the same 120s pass inheriting Npgsql's undocumented 30s default, and one of them was failing in production: on the dogfood box the `io_latency` baseline timed out nineteen times over two days and every sample landed between **30.1s and 31.4s**. A fixed wall, not a variable-duration fault and not the 15s connection timeout - the shipped query (dumped from the built assembly, not retyped) runs in **~1.6s** against the live store on the three busiest servers, so it crosses the ceiling only when the store stalls, which is why it read as intermittent. [#2820] had already made that query 5.6x faster (23.7s to 4.2s) and [#2826] had made the failure visible; neither touched what it was failing against, exactly as [#2827] and [#2826] had not touched [#2810]'s. The cost when it fires is larger than one pass: `GetOrComputeBaselinesAsync` assigns `_cache[cacheKey]` unconditionally, so a null result is CACHED and a single stall silences that (server, metric) pair's anomaly detection for the full 1-hour `CacheTtl`. **60s, bounded on both sides.** Below: the failure record is RIGHT-CENSORED - every run was killed at the ceiling, so nothing says whether it wanted 35s or 300s, and the value is therefore chosen as twice what it had rather than as a fitted number the data cannot support. Above - the bound that actually binds: `DarlingWorker.s_analysisTimeout` gives the pass 120s and the fact collector's thirty-one reads, up to eleven baseline computations per server, the anomaly detector and the drill-down all share it, so a per-command deadline near that figure would let ONE stalled command consume the pass and cost the server every other fact - strictly worse than the failure being fixed, which loses one metric. It matches `PgFactCollector.FactCommandTimeoutSeconds` deliberately, because two numbers in one budget would have to be reasoned about together every time either moved. Swept by SHAPE across the whole assembly rather than the one reported site (the [#2344] and `PgStatementText` mistake, twice paid for): `PgAnomalyDetector` 12, `PgFindingStore` 8, `PgDrillDownCollector` 10, `PgBaselineProvider` 1, `DarlingAnalysisService` 1. The drill-down's ten take that collector's OWN existing `DrillDownCommandTimeoutSeconds = 30` rather than a new value, so an author's deliberate choice on its other eight sites is not silently overridden. Pinned structurally over the pass - every command constructed in the assembly sets a deadline, and the value stays inside the justified band - and proven red three ways, each failing a DIFFERENT assertion: reverting the reported site, restoring 30s, and raising to the full 120s budget. Repo-wide this leaves 133 untimed `NpgsqlCommand` sites outside the analysis pass, which have different budgets and are filed separately rather than swept in here.
Expand Down Expand Up @@ -3206,4 +3207,5 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
[#2837]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2837
[#2870]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2870
[#2864]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2864
[#2874]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2874
[#2876]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2876
Loading
Loading