diff --git a/CHANGELOG.md b/CHANGELOG.md
index 7dc9131166..eca6197b9f 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -9,6 +9,12 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Fixed
+- **The dimension GC prunes against the oldest SURVIVING digest-carrying fact row, not an assumed horizon** ([#1813], closes [#1795]) - the proper closure of the #1782 x #1784 interaction those issues shipped with: query_stats and procedure_stats are both coverage-gated tiers AND dim-feeding tables, so on a coverage-lagging store the #1784 clamp held their purges every sweep, the #1782 guard read that as "purge did not complete" and deferred the whole dimension GC - every sweep, until a backfill landed. Proven by execution: a 400-day-old orphan dimension row survived with nothing failed anywhere. The deferral was logically correct (pruning content the held facts reference would dangle digests and the resolving view would serve NULL payload silently) but blunt: it stopped the GC entirely rather than narrowing it.
+
+ The true safety boundary is measured, not assumed: content older than the oldest surviving digest-carrying fact row cannot be referenced by anything, whatever the reason those facts are still there. The GC now probes each dim-feeding table's floor - `min(collection_time)` under exactly the predicate of a new V39 partial index per table, an index-edge read once per sweep rather than the oldest-chunk walk the pre-#1767 NULL-digest rows would force, the bounded shape the last_seen watermark design (#1768) demands - takes the minimum across tables (the same reasoning as the widest-retention horizon), and clamps the assumed cutoff to one day before it, the same margin the assumed side carries for the hourly last_seen refresh guard. On a healthy store the measured floor sits inside retention and the assumed horizon rules unchanged; on a held store the GC stays bounded instead of stopped, reclaiming content for queries that stopped running long before while everything the held facts reference survives. The blunt guard is thereby unnecessary rather than dormant - it survives in exactly one honest form: a floor that cannot be MEASURED (the table missing or unreachable) still defers the cycle, because pruning on an unknown boundary is the one way to dangle digests, and bounded growth beats silent corruption. The probe predicate, the V39 index predicate, and the dimension map are pinned to each other three ways, so a new dimension column lands in all of them or none. The viewer's schema ladder gains the matching V39 arm (index-existence sentinel, the V22 idiom) so a fully-migrated store maps to exactly the required version instead of tripping the connect gate.
+
+ Chasing the live proof surfaced a REAL test-fixture defect worth its own record: `EnsureContinuousAggregatesAsync` attaches refresh policies whose jobs fire immediately (the #1788 finding), and the class's deep force-refresh collided with them (55P03) while a restore's DROP could collide with a running job's lock and silently strand `query_stats_db_hourly`/`db_daily` in the shared fixture - stranded aggregates whose deep coverage then flipped the #1784 gate for every LATER test, the same manufactured-flake mechanism #1794 documented for connection debris. The tests never needed the scheduler (they refresh manually): the ensure wrapper now removes every rollup's refresh policy immediately, and the one force-refresh that can still catch an already-executing job retries bounded on 55P03 only - the product's own #1788 idiom. Three consecutive full-class live runs leave zero stranded aggregates.
+
- **Recommendations warm-up now survives the 512 MB archive/reset, and analysis reads the archive tier everywhere** ([#1811], closes [#1809]) - the 24-hour sufficiency check measured MIN..MAX(collection_time) on the RAW hot wait_stats table, while `ArchiveAllAndResetAsync` copies everything to Parquet and empties the hot store. On a multi-server install that trips the size threshold more often than daily, the measured span restarted with every reset and warm-up never completed - even though the archived history was sitting in Parquet and every other tab could read it through the `v_` union views. The check now reads `v_wait_stats`, so archived plus hot history counts, and a genuinely young install still gates (both pinned, watched red by reverting the query to the raw table).
The subtlety that made this a 24-file-line sweep rather than a one-word fix: the broken gate was accidentally SHIELDING a second instance of the same defect. The analysis fact collector read raw tables in 22 more places - the window reads (query stats, snapshots, memory, perfmon, file IO, blocking, deadlocks, storage) and the on-load config snapshots (server/database config, trace flags, server properties), which after a reset are EMPTY until the next app start. Fixing only the gate would have run analysis over a thin post-reset hot window, reintroducing the exact fraction-of-period distortion the 24-hour floor exists to prevent ("5 seconds of THREADPOOL looks alarming in a 16-minute window"). All 23 sites now read their `v_` archive views - column-identical by construction (only `config_alert_log` carries a view-only column, and analysis does not read it), dedup-safe where re-collection could duplicate rows (the QUALIFY views), and a catalog-driven sweep guard holds the whole pipeline there: for every table in `ArchiveService.ArchivableTables`, a raw `FROM` anywhere in Lite/Analysis fails the build's tests naming the file - so a new collector's table is guarded the day it exists.
@@ -1902,6 +1908,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
[#1803]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/1803
[#1805]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/1805
[#1808]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/1808
+[#1795]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/1795
+[#1813]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/1813
[#1809]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/1809
[#1811]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/1811
[#1794]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/1794
diff --git a/Darling/Darling.Tests/DarlingDimensionGcBoundTests.cs b/Darling/Darling.Tests/DarlingDimensionGcBoundTests.cs
new file mode 100644
index 0000000000..f4005b409b
--- /dev/null
+++ b/Darling/Darling.Tests/DarlingDimensionGcBoundTests.cs
@@ -0,0 +1,87 @@
+/*
+ * Copyright (c) 2026 Erik Darling, Darling Data LLC
+ *
+ * This file is part of the SQL Server Performance Monitor.
+ *
+ * Licensed under the MIT License. See LICENSE file in the project root for full license information.
+ */
+
+using System;
+using System.Linq;
+using PerformanceMonitor.Darling.Service;
+using PerformanceMonitor.Darling.Storage;
+using Xunit;
+
+namespace Darling.Tests;
+
+///
+/// #1795: the dimension GC's cutoff math and the three-way alignment behind its floor probe. The probe's
+/// speed contract is that its WHERE clause EXACTLY matches the V39 partial index predicate, and both are
+/// derived from — these pins hold the derived strings AND the V39 DDL
+/// to each other, so a new dimension column cannot land in the map without landing in the index, and the
+/// index predicate cannot drift from the probe's.
+///
+public sealed class DarlingDimensionGcBoundTests
+{
+ private static readonly DateTime Now = new(2026, 7, 28, 12, 0, 0, DateTimeKind.Unspecified);
+
+ /* widest = 30 → assumed cutoff = now - (30 + ChunkIntervalDays + 1). */
+ private static DateTime Assumed => Now.AddDays(-(30 + TimescaleSupport.ChunkIntervalDays + 1));
+
+ [Fact]
+ public void HealthyFloor_LeavesTheAssumedHorizonAlone()
+ {
+ /* Floor well inside retention (facts purging normally): the measured bound (floor - 1d) sits
+ NEWER than the assumed cutoff, and the cutoff must not move forward past the assumed horizon —
+ the GC never gets MORE aggressive than today. */
+ var cutoff = DarlingRetention.ComputeDimensionCutoff(Now, 30, Now.AddDays(-4));
+ Assert.Equal(Assumed, cutoff);
+ }
+
+ [Fact]
+ public void HeldFloor_ClampsTheCutoffToOneDayBeforeIt()
+ {
+ /* The #1795 field state: the clamp holds 45-day-old facts, older than the assumed horizon. The
+ cutoff follows the MEASURED floor minus the one-day last_seen margin, so content those facts
+ reference survives while anything older is reclaimed. */
+ var floor = Now.AddDays(-45);
+ var cutoff = DarlingRetention.ComputeDimensionCutoff(Now, 30, floor);
+ Assert.Equal(floor.AddDays(-1), cutoff);
+ Assert.True(cutoff < Assumed);
+ }
+
+ [Fact]
+ public void NoDigestFacts_FallBackToTheAssumedHorizon()
+ {
+ /* A fresh (or fully-aged) store has no digest-carrying facts at all: nothing can dangle, and
+ last_seen still bounds what is old enough to take. */
+ var cutoff = DarlingRetention.ComputeDimensionCutoff(Now, 30, oldestSurvivingDigestFact: null);
+ Assert.Equal(Assumed, cutoff);
+ }
+
+ [Fact]
+ public void DigestPredicates_AreExactlyTheDeclaredColumns_InDeclarationOrder()
+ {
+ /* The strings themselves, pinned: the probe filters on these, the V39 index is declared with
+ these, and both derive from PayloadDimensions.All. */
+ Assert.Equal(
+ "query_text_digest IS NOT NULL OR query_plan_digest IS NOT NULL",
+ PayloadDimensions.DigestPredicateByTable["query_stats"]);
+ Assert.Equal(
+ "query_plan_digest IS NOT NULL",
+ PayloadDimensions.DigestPredicateByTable["procedure_stats"]);
+ Assert.Equal(2, PayloadDimensions.DigestPredicateByTable.Count);
+ }
+
+ [Fact]
+ public void V39Indexes_UseExactlyTheProbePredicates()
+ {
+ var v39 = PgMigrations.Scripts.Single(m => m.Version == 39).Sql;
+
+ foreach (var (factTable, predicate) in PayloadDimensions.DigestPredicateByTable)
+ {
+ Assert.Contains($"ON {factTable} (collection_time)", v39, StringComparison.Ordinal);
+ Assert.Contains($"WHERE {predicate}", v39, StringComparison.Ordinal);
+ }
+ }
+}
diff --git a/Darling/Darling.Tests/DarlingObservabilityTests.cs b/Darling/Darling.Tests/DarlingObservabilityTests.cs
index 6e8fe288e3..7c155ad5f0 100644
--- a/Darling/Darling.Tests/DarlingObservabilityTests.cs
+++ b/Darling/Darling.Tests/DarlingObservabilityTests.cs
@@ -76,8 +76,8 @@ public void MigrationScripts_AreRegisteredInAscendingOrder_V34AgCollectors_V36Ag
Assert.Equal(33, PgMigrations.Scripts[32].Version);
/* The newest migration is asserted by identity rather than by ordinal: this ladder is walked by every
stacked branch at once, and a positional pin turns each addition into a conflict for the next. */
- Assert.Equal(38, PgMigrations.Scripts[^1].Version);
- Assert.Equal(38, StorageVersion.SchemaVersion);
+ Assert.Equal(39, PgMigrations.Scripts[^1].Version);
+ Assert.Equal(39, StorageVersion.SchemaVersion);
/* V34 (#991) creates the two Availability Group collector tables. Schema-qualified collect.* and
CREATE TABLE IF NOT EXISTS, per the file's additive-create idiom (V29): a no-op on a fresh store
diff --git a/Darling/Darling.Tests/PayloadDimensionLiveTests.cs b/Darling/Darling.Tests/PayloadDimensionLiveTests.cs
index a96605c20a..d99b3873fb 100644
--- a/Darling/Darling.Tests/PayloadDimensionLiveTests.cs
+++ b/Darling/Darling.Tests/PayloadDimensionLiveTests.cs
@@ -56,10 +56,20 @@ public sealed class PayloadDimensionLiveTests
/// The exact line the dimension GC emits when it stands down, pinned in FULL rather than by prefix: it is
/// a field signature an operator greps for, and a partial pin would let the wording drift to name a cause
/// the guard does not actually detect -- which is how it came to say "raw purges held" while the shipped
- /// trigger was a purge that did not complete.
+ /// trigger was a purge that did not complete. Since #1795 the ONLY deferral trigger is an UNMEASURABLE
+ /// fact floor (the measured bound replaced the blunt purge-failed guard), and the line says so.
///
private const string DeferralSignature =
- "dimension GC deferred: a dim-feeding table kept rows past its horizon this cycle; dimension content is retained until its purge completes";
+ "dimension GC deferred: a dim-feeding table's fact floor was unmeasurable this cycle; dimension content is retained until it can be measured";
+
+ ///
+ /// The exact line the GC emits when the MEASURED bound clamps below the assumed horizon (#1795) —
+ /// the greppable signature of a store whose dimension GC is bounded by held fact history (a coverage
+ /// clamp or a failed purge) rather than by the nominal horizon. Informational: this state is the fix
+ /// working.
+ ///
+ private const string BoundedSignature =
+ "dimension GC bounded by surviving facts: dimension content newer than the oldest digest-carrying fact row is retained";
private const string SkipReason =
"Set DARLING_TEST_PG to a Postgres connection string to run the payload-dimension store tests.";
@@ -795,21 +805,22 @@ await LiveStoreCleanup.RunAsync(connectionString!, bodySucceeded, async (cleanup
}
///
- /// The dimension GC must NOT prune while a table that feeds it failed to purge this cycle (#1782).
+ /// The dimension GC must NOT prune while a dim-feeding table's fact floor is UNMEASURABLE (#1795,
+ /// narrowing #1782's guard).
///
- /// The cutoff assumes facts past their retention are gone. The fact loop is failure-isolated, so a
- /// table whose drop_chunks fails AND whose DELETE fallback then also fails is warned, skipped, and the
- /// sweep continues straight into the GC — pruning dimension rows those surviving facts still reference.
- /// The digests dangle and the resolving view serves NULL payload with no error raised anywhere, which is
- /// the whole reason this is a guard and not a comment.
+ /// The GC's cutoff is bounded by the oldest surviving digest-carrying fact row. When that floor
+ /// cannot be measured at all, the safety boundary is unknown — pruning on the assumed horizon could
+ /// dangle digests that still-unmeasurable facts reference, and the resolving view would serve NULL
+ /// payload with no error raised anywhere. So an unmeasurable floor defers the GC for the cycle, the
+ /// same fail-toward-recoverable choice the old purge-failed guard made.
///
- /// The failure is injected by RENAMING query_stats for the duration, which is a genuine total
- /// failure of both purge paths (the relation the statements name does not exist) rather than a simulated
- /// one — and it exercises the real seam: the sweep sets the flag, and the GC reads it. It is restored in
- /// a finally, and the whole live collection is serialized, so no sibling class can observe the rename.
+ /// The unmeasurable state is injected by RENAMING query_stats for the duration — the floor probe's
+ /// relation genuinely does not exist, exercising the real seam rather than a simulated flag. It is
+ /// restored in a finally, and the whole live collection is serialized, so no sibling class can observe
+ /// the rename.
///
[Fact]
- public async Task DimensionGc_DefersWhenADimFeedingTableFailedItsPurge_ThenPrunesOnceItSucceeds()
+ public async Task DimensionGc_DefersWhenAFactFloorIsUnmeasurable_ThenPrunesOnceItIs()
{
var connectionString = Environment.GetEnvironmentVariable("DARLING_TEST_PG");
Assert.SkipWhen(string.IsNullOrEmpty(connectionString), SkipReason);
@@ -866,14 +877,15 @@ must land on. Asserting the log line first would let a broken guard go red on mi
renamed = false;
}
- /* Since #1784 the sweep SKIPS a raw table whose rollup does not cover it, and that skip sets the
- same dim-feeding-purge-failed flag this test is about — so a "healthy purge" now means covered as
- well as succeeding. Satisfy coverage before the control phase, or the GC defers for the OTHER
- reason and the control proves nothing. */
+ /* Coverage satisfied so the control phase exercises the fully-healthy path end to end — no
+ clamp skip in the log, floors measurable, purge running. (A clamp skip would no longer
+ defer the GC — that is #1795's whole point, pinned by the sibling test below — but the
+ control here is about the unmeasurable-floor deferral ENDING.) */
await SatisfyRollupCoverageAsync(connection, ct);
- /* Control: with the purge healthy again the guard opens and the same row is pruned, so the test
- cannot pass merely because the GC never runs. */
+ /* Control: with the table back the floor is measurable again. No digest-carrying facts exist
+ on this rig (the floor is NULL), so the assumed ~32-day horizon applies and the 400-day
+ orphan is pruned — the test cannot pass merely because the GC never runs. */
var prunedLog = new CapturingTestLogger();
await DarlingRetention.PurgeAsync(postgres, timescaleAvailable: true, prunedLog, ct, PurgeDimFeeding(30));
@@ -897,9 +909,12 @@ await LiveStoreCleanup.RunAsync(connectionString!, bodySucceeded, async (cleanup
await restore.ExecuteNonQueryAsync(cleanupCt);
}
- await RestoreCaggsAsync(cleanup, coverageSnapshot, cleanupCt);
-
await DeleteDimRowAsync(cleanup, PayloadDimensions.QueryPlanDimTable, digest, cleanupCt);
+
+ /* LAST, deliberately: the CAGG drops are the one step that can break the session (a
+ drop colliding with a still-running background job — the CI-proven shape), and a
+ break here has no cleanup statements left to strand. */
+ await RestoreCaggsAsync(cleanup, coverageSnapshot, cleanupCt);
});
}
}
@@ -912,6 +927,124 @@ private static async Task DimRowCountAsync(NpgsqlConnection connection, by
return (long)(await read.ExecuteScalarAsync(ct))!;
}
+ ///
+ /// The #1795 property, in the exact field shape that motivated it: on a coverage-lagging store the
+ /// #1784 clamp holds query_stats' drop EVERY sweep, and before this change that set the #1782 defer
+ /// flag — so the dimension GC stood down every cycle and a 400-day orphan survived with nothing
+ /// failed anywhere. The GC now prunes against the oldest SURVIVING digest-carrying fact row: the
+ /// orphan (older than anything still referencing content) is reclaimed even while the clamp holds,
+ /// and content the held facts still reference survives, because pruning it would dangle digests the
+ /// clamp deliberately kept alive.
+ ///
+ [Fact]
+ public async Task DimensionGc_PrunesOrphansButKeepsReferencedContent_WhileTheCoverageClampHoldsFacts()
+ {
+ var connectionString = Environment.GetEnvironmentVariable("DARLING_TEST_PG");
+ Assert.SkipWhen(string.IsNullOrEmpty(connectionString), SkipReason);
+
+ var ct = TestContext.Current.CancellationToken;
+ await using var connection = await OpenMigratedStoreAsync(connectionString!, ct);
+
+ Assert.True(await TimescaleSupport.TryEnableAsync(connection, null, ct), "the rig is expected to have TimescaleDB");
+ await TimescaleSupport.ConvertToHypertablesAsync(connection, null, ct);
+
+ var preexistingCaggs = await ExistingCaggsAsync(connection, ct);
+
+ var (serverId, serverName) = NewServer();
+
+ /* The surviving fact: 45 days back — past the 30-day catalog horizon (the clamp is what keeps
+ it alive) and past the ~32-day assumed GC cutoff, so only the MEASURED bound protects the
+ content it references. */
+ var heldFactTime = DateTime.SpecifyKind(DateTime.UtcNow.AddDays(-45), DateTimeKind.Unspecified);
+ var referencedDigest = PayloadDimensions.Digest("pm1795-referenced-" + Guid.NewGuid().ToString("N"));
+ var orphanDigest = PayloadDimensions.Digest("pm1795-orphan-" + Guid.NewGuid().ToString("N"));
+ var orphanSeen = DateTime.SpecifyKind(DateTime.UtcNow.AddDays(-400), DateTimeKind.Unspecified);
+
+ await using var postgres = NpgsqlDataSource.Create(connectionString!);
+
+ var bodySucceeded = false;
+ try
+ {
+ /* INSIDE the try: creating aggregates is the shared-fixture mutation the finally restores.
+ Policies removed immediately - see EnsureAggregatesWithoutPoliciesAsync. */
+ await EnsureAggregatesWithoutPoliciesAsync(connection, ct);
+
+ /* The held digest-carrying fact... */
+ using (var seed = new NpgsqlCommand(
+ "INSERT INTO query_stats (collection_id, collection_time, server_id, server_name, database_name, query_hash, query_plan_digest) " +
+ "VALUES ($1, $2, $3, $4, 'pm1795', '0xPM1795', $5)", connection))
+ {
+ seed.Parameters.AddWithValue(CollectionIdGenerator.Next());
+ seed.Parameters.AddWithValue(heldFactTime);
+ seed.Parameters.AddWithValue(serverId);
+ seed.Parameters.AddWithValue(serverName);
+ seed.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Bytea, Value = referencedDigest });
+ await seed.ExecuteNonQueryAsync(ct);
+ }
+
+ /* ...the content it references (last_seen = its sighting time, 45d — inside the measured
+ bound of floor minus one day)... */
+ using (var seed = new NpgsqlCommand(
+ $"INSERT INTO {PayloadDimensions.QueryPlanDimTable} (digest, query_plan_xml, last_seen) " +
+ "VALUES ($1, $2, $3) ON CONFLICT (digest) DO NOTHING", connection))
+ {
+ seed.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Bytea, Value = referencedDigest });
+ seed.Parameters.AddWithValue("");
+ seed.Parameters.AddWithValue(heldFactTime);
+ await seed.ExecuteNonQueryAsync(ct);
+ }
+
+ /* ...and the orphan nothing references, older than the oldest surviving fact. */
+ using (var seed = new NpgsqlCommand(
+ $"INSERT INTO {PayloadDimensions.QueryPlanDimTable} (digest, query_plan_xml, last_seen) " +
+ "VALUES ($1, $2, $3) ON CONFLICT (digest) DO NOTHING", connection))
+ {
+ seed.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Bytea, Value = orphanDigest });
+ seed.Parameters.AddWithValue("");
+ seed.Parameters.AddWithValue(orphanSeen);
+ await seed.ExecuteNonQueryAsync(ct);
+ }
+
+ /* Coverage LAGS: the aggregate holds recent buckets only, none reaching the 45-day fact, so
+ the clamp holds query_stats' drop — the field state. */
+ using (var refresh = new NpgsqlCommand(
+ TimescaleSupport.RefreshContinuousAggregateSql(TimescaleSupport.QueryStatsHourlyView), connection))
+ {
+ refresh.Parameters.AddWithValue(DateTime.SpecifyKind(DateTime.UtcNow.AddDays(-2), DateTimeKind.Unspecified));
+ await refresh.ExecuteNonQueryAsync(ct);
+ }
+
+ Assert.False(
+ await TimescaleSupport.IsRawTierDropSafeAsync(connection, "query_stats", ct),
+ "the clamp must be holding query_stats for this test to mean anything");
+
+ var log = new CapturingTestLogger();
+ await DarlingRetention.PurgeAsync(postgres, timescaleAvailable: true, log, ct, PurgeDimFeeding(30));
+
+ /* The clamp held (the fact survived its horizon) — and the GC ran ANYWAY, bounded. */
+ Assert.Contains(CoverageSkipSignature, log.Joined, StringComparison.Ordinal);
+ Assert.Equal(1L, await AncientRowCountAsync(connection, serverId, ct));
+ Assert.DoesNotContain(DeferralSignature, log.Joined, StringComparison.Ordinal);
+ Assert.Contains(BoundedSignature, log.Joined, StringComparison.Ordinal);
+
+ /* The substantive pair: the orphan is reclaimed, the referenced content is not. */
+ Assert.Equal(0L, await DimRowCountAsync(connection, orphanDigest, ct));
+ Assert.Equal(1L, await DimRowCountAsync(connection, referencedDigest, ct));
+
+ bodySucceeded = true;
+ }
+ finally
+ {
+ await LiveStoreCleanup.RunAsync(connectionString!, bodySucceeded, async (cleanup, cleanupCt) =>
+ {
+ await DeleteServerRowsAsync(cleanup, serverId, cleanupCt);
+ await DeleteDimRowAsync(cleanup, PayloadDimensions.QueryPlanDimTable, referencedDigest, cleanupCt);
+ await DeleteDimRowAsync(cleanup, PayloadDimensions.QueryPlanDimTable, orphanDigest, cleanupCt);
+ await RestoreCaggsAsync(cleanup, preexistingCaggs, cleanupCt);
+ });
+ }
+ }
+
///
/// The catalog sweep must not drop raw chunks the rollup has not captured (#1784).
///
@@ -951,8 +1084,9 @@ test adds — see ExistingCaggsAsync. */
try
{
/* INSIDE the try: creating aggregates is the shared-fixture mutation the finally restores,
- so it must not happen where a throw would skip that restore. */
- await TimescaleSupport.EnsureContinuousAggregatesAsync(connection, null, ct);
+ so it must not happen where a throw would skip that restore. Policies removed immediately -
+ their instantly-firing jobs are the 55P03/stranded-aggregate race this class chased in #1795. */
+ await EnsureAggregatesWithoutPoliciesAsync(connection, ct);
using (var seed = new NpgsqlCommand(
"INSERT INTO query_stats (collection_id, collection_time, server_id, server_name, database_name, query_hash) " +
@@ -1042,12 +1176,45 @@ private static async Task TryExecAsync(NpgsqlConnection connection, string sql,
{
try
{
+ /* A PRIOR best-effort statement may have swallowed a failure that BROKE this connection
+ (proven on CI: a RestoreCaggs drop died, the swallow hid it, and the next helper threw
+ "Connection is not open" out of an otherwise-green test). Reopening the same
+ NpgsqlConnection object checks out a fresh pooled session, so one broken statement
+ cannot cascade into every cleanup statement after it. */
+ if (connection.State != System.Data.ConnectionState.Open)
+ {
+ await connection.OpenAsync(ct);
+ }
+
using var command = new NpgsqlCommand(sql, connection);
await command.ExecuteNonQueryAsync(ct);
}
- catch (Exception)
+ catch (Exception ex)
+ {
+ /* Cleanup is best-effort: a restore step that throws must not mask the test's own failure.
+ But say WHAT was swallowed — xUnit captures console per test, so on the next mystery this
+ line is the diagnosis instead of a silent hole. */
+ Console.WriteLine($"[cleanup best-effort] {ex.GetType().Name} ({(ex as PostgresException)?.SqlState}): {ex.Message} — {sql}");
+ }
+ }
+
+ ///
+ /// and then REMOVES every rollup's
+ /// refresh policy. The ensure attaches policies whose jobs fire IMMEDIATELY (the #1788 finding), and
+ /// those background refreshes race everything this class does next: a manual force-refresh collides
+ /// with 55P03 "concurrent refresh", and a restore's DROP can collide with a running job's lock and
+ /// silently strand the aggregate in the shared fixture — which is exactly how query_stats_db_hourly
+ /// and query_stats_db_daily came to persist across runs and flip the #1784 gate for every later
+ /// test. The tests refresh manually and deterministically; they never need the scheduler.
+ ///
+ private static async Task EnsureAggregatesWithoutPoliciesAsync(NpgsqlConnection connection, CancellationToken ct)
+ {
+ await TimescaleSupport.EnsureContinuousAggregatesAsync(connection, null, ct);
+
+ foreach (var (view, _, _, _) in TimescaleSupport.RollupViews)
{
- /* Cleanup is best-effort: a restore step that throws must not mask the test's own failure. */
+ await TryExecAsync(connection,
+ $"SELECT remove_continuous_aggregate_policy('collect.{view}', if_exists => true)", ct);
}
}
@@ -1059,18 +1226,34 @@ private static async Task TryExecAsync(NpgsqlConnection connection, string sql,
/// outside its try, so a throw part-way through creation still leaves the finally a baseline to restore
/// against. Owning the snapshot here made the restore conditional on this method RETURNING, which is
/// exactly the window in which it leaks aggregates into the shared fixture.
+ ///
+ /// The refresh retries bounded on 55P03: removing the policies closes the standing race, but a
+ /// job that was ALREADY EXECUTING when its policy was removed finishes its run, and one collision can
+ /// still land. The product's backfill slices carry the same bounded transient-only retry for the same
+ /// reason (#1788); every other SQLSTATE still fails fast.
///
private static async Task SatisfyRollupCoverageAsync(NpgsqlConnection connection, CancellationToken ct)
{
- await TimescaleSupport.EnsureContinuousAggregatesAsync(connection, null, ct);
+ await EnsureAggregatesWithoutPoliciesAsync(connection, ct);
foreach (var (relation, _, coverage) in TimescaleSupport.RawTierCoverage)
{
_ = relation;
- using var refresh = new NpgsqlCommand(
- TimescaleSupport.RefreshContinuousAggregateSql(coverage, force: true), connection);
- refresh.Parameters.AddWithValue(DateTime.SpecifyKind(DateTime.UtcNow.AddDays(-3650), DateTimeKind.Unspecified));
- await refresh.ExecuteNonQueryAsync(ct);
+ for (var attempt = 1; ; attempt++)
+ {
+ try
+ {
+ using var refresh = new NpgsqlCommand(
+ TimescaleSupport.RefreshContinuousAggregateSql(coverage, force: true), connection);
+ refresh.Parameters.AddWithValue(DateTime.SpecifyKind(DateTime.UtcNow.AddDays(-3650), DateTimeKind.Unspecified));
+ await refresh.ExecuteNonQueryAsync(ct);
+ break;
+ }
+ catch (PostgresException ex) when (ex.SqlState == "55P03" && attempt < 5)
+ {
+ await Task.Delay(TimeSpan.FromMilliseconds(250 * attempt), ct);
+ }
+ }
}
}
diff --git a/Darling/Darling.Tests/ViewerDataServiceTests.cs b/Darling/Darling.Tests/ViewerDataServiceTests.cs
index a5b0c74a8f..70aaf91e6c 100644
--- a/Darling/Darling.Tests/ViewerDataServiceTests.cs
+++ b/Darling/Darling.Tests/ViewerDataServiceTests.cs
@@ -479,7 +479,7 @@ public void RequiredStoreSchemaVersion_TracksTheBuildSchemaVersion_AndTheProbeCo
the connect-time gate refuse to open the viewer against a perfectly healthy store. */
Assert.Equal(
ViewerDataService.RequiredStoreSchemaVersion,
- ViewerDataService.MapProbedSchemaVersion(true, true, true, true, true, true, true, true, true, true, true, true, true, true, true, true, true, true, true, true, true, true));
+ ViewerDataService.MapProbedSchemaVersion(true, true, true, true, true, true, true, true, true, true, true, true, true, true, true, true, true, true, true, true, true, true, true));
}
}
diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingRetention.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingRetention.cs
index a9853c148b..22f30a5f7c 100644
--- a/Darling/PerformanceMonitor.Darling.Service/DarlingRetention.cs
+++ b/Darling/PerformanceMonitor.Darling.Service/DarlingRetention.cs
@@ -140,12 +140,6 @@ public static async Task PurgeAsync(
/* Naive-UTC storage: Npgsql 6+ rejects Kind=Utc against `timestamp` — see PgCollectorRowWriter. */
var utcNow = DateTime.SpecifyKind(DateTime.UtcNow, DateTimeKind.Unspecified);
- /* Set when a table that feeds a payload dimension could not be purged this cycle, which invalidates
- the dimension GC's cutoff (#1782). Tracked here rather than inferred later because the sweep is the
- only thing that knows: a failed purge is warned and skipped, leaving no state a later query could
- read back. */
- var dimFeedingPurgeFailed = false;
-
try
{
foreach (var definition in CollectorCatalog.All)
@@ -157,7 +151,6 @@ drift degrades to a loud warning (and a WARNING run-record) instead of killing t
logger?.LogWarning("Retention purge: no schedule entry for '{Collector}' — {Table} was not purged",
definition.Name, definition.TargetTable);
tablesFailed++;
- dimFeedingPurgeFailed |= PayloadDimensions.ForTable(definition.TargetTable).Count > 0;
continue;
}
@@ -184,9 +177,9 @@ So the same predicate now guards both. It is binary rather than a clamped cutoff
"Retention purge SKIPPED for {Table}: its rollup does not yet cover the oldest rows, so dropping would delete history no aggregate holds. Resumes by itself once a backfill extends coverage.",
definition.TargetTable);
- /* Those rows are still there past their horizon, which is the dimension GC's hazard
- (#1782) exactly as a failed purge is — the GC's cutoff assumes they are gone. */
- dimFeedingPurgeFailed |= PayloadDimensions.ForTable(definition.TargetTable).Count > 0;
+ /* Those rows are still there past their horizon — which is fine for the dimension GC
+ below (#1795): its cutoff is MEASURED from the oldest surviving digest-carrying fact
+ row, so held history bounds the GC instead of deferring it. */
continue;
}
@@ -219,10 +212,9 @@ table still honors its horizon instead of growing unbounded. */
{
tablesFailed++;
- /* Both paths failed for this table, so its rows are still there past their horizon. If it
- feeds a payload dimension, that invalidates the dimension GC's cutoff for this cycle
- (#1782) — see the guard below. */
- dimFeedingPurgeFailed |= PayloadDimensions.ForTable(definition.TargetTable).Count > 0;
+ /* Both paths failed for this table, so its rows are still there past their horizon.
+ The dimension GC below stays safe anyway (#1795): its cutoff is measured from the
+ oldest surviving digest-carrying fact row, which those rows ARE. */
}
}
@@ -260,40 +252,67 @@ the dims and orphan a reader — plus a margin covering the two ways a fact can
widestFactRetentionDays = Math.Max(widestFactRetentionDays, factRetentionDays);
}
- /* That whole cutoff embeds one assumption: facts older than their retention are GONE. It holds
- while the fact purge SUCCEEDS — and the loop above is failure-isolated, so a table whose
- drop_chunks fails AND whose DELETE fallback then also fails is warned, skipped, and the sweep
- continues straight into this GC. Those facts persist past the cutoff while their dimension rows
- are pruned on schedule: the digests dangle, and the resolving view serves NULL payload with no
- error raised anywhere (#1782).
-
- So a dim-feeding table failing its purge defers the GC for that cycle. It is keyed to the
- failure rather than to the retention POLICY's held state, because a held policy cannot cause
- this: the catalog sweep above drops these tables at their own 30-day horizon regardless of the
- policy, two days inside this 32-day cutoff, and the ordering survives any override since both
- sides resolve through the same retentionDaysFor. Deferring on a held policy would stall the GC
- indefinitely on a store whose policies are held for entirely unrelated reasons.
-
- Self-ending by construction: the next sweep that purges cleanly runs the GC, and nothing needs
- an operator. And it fails toward the recoverable side — deferring costs bounded dimension
- growth, visible and reclaimed by the next cycle, where pruning early corrupts the read path
- silently. */
- if (dimFeedingPurgeFailed)
+ /* That assumed cutoff embeds one assumption: facts older than their retention are GONE. Two
+ legitimate mechanisms break it — a dim-feeding table whose purge FAILED (#1782), and one the
+ coverage clamp above deliberately SKIPPED (#1784). The old response deferred the whole GC
+ whenever either happened, which was correct but blunt: on a coverage-lagging store the clamp
+ holds every sweep until a backfill lands, so the GC deferred every sweep and a 400-day-old
+ orphan survived with nothing failed anywhere (#1795).
+
+ The true safety boundary is MEASURED instead: content older than the oldest SURVIVING
+ digest-carrying fact row cannot be referenced by anything, whatever the reason those facts
+ are still there. Probe each dim-feeding table's floor — an index-edge read through its V39
+ partial index, once per sweep, the bounded shape the last_seen watermark design (#1768)
+ demands — take the MINIMUM across tables (same reasoning as the widest retention above), and
+ clamp the assumed cutoff to it. The blunt guard becomes unnecessary rather than dormant: a
+ store whose purges are blocked for weeks still reclaims dimension content for queries that
+ stopped running long before, and held history bounds the GC instead of stopping it.
+
+ A floor that cannot be MEASURED (the table missing/renamed/unreachable) is different: the
+ safety boundary is then unknown, so the GC defers for the cycle — fail toward the recoverable
+ side, exactly as the old guard did. Bounded growth beats silently dangled digests. */
+ DateTime? oldestSurvivingDigestFact = null;
+ var floorUnmeasurable = false;
+ foreach (var (factTable, digestPredicate) in PayloadDimensions.DigestPredicateByTable)
+ {
+ try
+ {
+ await using var floorConnection = await postgres.OpenConnectionAsync(cancellationToken);
+ await using var probe = new NpgsqlCommand(
+ $"SELECT min({PrefixTimeColumn(factTable)}) FROM {factTable} WHERE {digestPredicate}", floorConnection);
+ var floor = await probe.ExecuteScalarAsync(cancellationToken);
+ if (floor is DateTime floorTime
+ && (oldestSurvivingDigestFact is null || floorTime < oldestSurvivingDigestFact.Value))
+ {
+ oldestSurvivingDigestFact = floorTime;
+ }
+ }
+ catch (Exception ex) when (ex is not OperationCanceledException)
+ {
+ floorUnmeasurable = true;
+ logger?.LogWarning("Retention purge: could not measure {Table}'s digest-carrying fact floor: {Message}",
+ factTable, ex.Message);
+ }
+ }
+
+ if (floorUnmeasurable)
{
- /* One line, deliberately a fixed string never interpolated: it is the field signature an
- operator greps for when the dimensions stop shrinking, so it must not vary by store. It
- names the observable STATE -- a table still holding rows past its horizon -- and not any
- cause of it. Two different things produce that state: a purge statement that failed, and a
- purge deliberately skipped because its rollup does not cover the table (#1784). Naming
- either one here would be a lie in the other case and would send the reader after the wrong
- remedy, which is the one thing a diagnostic line must not do. The preceding per-table
- warning already says which table and which of the two it was, so the cause is one line
- above, owned by the code that actually knows it. */
- logger?.LogWarning("dimension GC deferred: a dim-feeding table kept rows past its horizon this cycle; dimension content is retained until its purge completes");
+ /* One line, deliberately a fixed string never interpolated: the field signature an operator
+ greps for when the dimensions stop shrinking. The per-table warning above names which
+ table and why. */
+ logger?.LogWarning("dimension GC deferred: a dim-feeding table's fact floor was unmeasurable this cycle; dimension content is retained until it can be measured");
}
else
{
- var dimensionCutoff = utcNow.AddDays(-(widestFactRetentionDays + TimescaleSupport.ChunkIntervalDays + 1));
+ var dimensionCutoff = ComputeDimensionCutoff(utcNow, widestFactRetentionDays, oldestSurvivingDigestFact);
+ if (dimensionCutoff < utcNow.AddDays(-(widestFactRetentionDays + TimescaleSupport.ChunkIntervalDays + 1)))
+ {
+ /* Fixed string, same reasoning as the defer line: the greppable signature of a store
+ whose GC is bounded by held history (clamped or failed purges) rather than by the
+ nominal horizon. Informational — this state is the fix working, not a problem. */
+ logger?.LogInformation("dimension GC bounded by surviving facts: dimension content newer than the oldest digest-carrying fact row is retained");
+ }
+
foreach (var dimTable in PayloadDimensions.DimTables)
{
var dimDeleted = await PurgeOneAsync(
@@ -494,6 +513,37 @@ internal static string TimeSlicedDeleteSql(string table, string timeColumn, stri
+ $" AND {timeColumn} < (SELECT min({timeColumn}) FROM {table} WHERE {expired}) + INTERVAL '{TimescaleSupport.ChunkIntervalDays} days'";
}
+ ///
+ /// The dimension GC's cutoff (#1795): the ASSUMED horizon (widest dim-feeding fact retention +
+ /// drop_chunks granularity + 1 day for the hourly
+ /// last_seen refresh guard), CLAMPED to one day before the oldest surviving digest-carrying
+ /// fact row when that measured floor reaches further back — held history bounds the GC instead of
+ /// deferring it. The measured side carries the SAME one-day margin, for the same reason: a dim row's
+ /// last_seen can trail its newest referencing fact by up to the hourly refresh guard, so
+ /// pruning right AT the floor could take content the floor row still references. A null floor (no
+ /// digest-carrying facts anywhere — a fresh or fully-aged store) leaves the assumed horizon alone:
+ /// with no facts, nothing can dangle, and last_seen still bounds what is old enough to take.
+ ///
+ internal static DateTime ComputeDimensionCutoff(DateTime utcNow, int widestFactRetentionDays, DateTime? oldestSurvivingDigestFact)
+ {
+ var assumed = utcNow.AddDays(-(widestFactRetentionDays + TimescaleSupport.ChunkIntervalDays + 1));
+ if (oldestSurvivingDigestFact is null)
+ {
+ return assumed;
+ }
+
+ var measured = oldestSurvivingDigestFact.Value.AddDays(-1);
+ return measured < assumed ? measured : assumed;
+ }
+
+ ///
+ /// The prefix time column of a dim-feeding fact table, resolved from the catalog so the floor probe
+ /// can never disagree with the table's actual schema (both current dim-feeding tables use
+ /// collection_time; the resolver keeps that true by construction rather than by assertion).
+ ///
+ private static string PrefixTimeColumn(string factTable) =>
+ CollectorCatalog.All.First(c => string.Equals(c.TargetTable, factTable, StringComparison.Ordinal)).PrefixTimeColumnName;
+
///
/// The Timescale purge statement for one collector table — drop_chunks detaches every
/// chunk wholly older than the horizon (validated live on TimescaleDB 2.28.1; the partition
diff --git a/Darling/PerformanceMonitor.Darling.Storage/PayloadDimensions.cs b/Darling/PerformanceMonitor.Darling.Storage/PayloadDimensions.cs
index cbc33f6dbd..62592f93f1 100644
--- a/Darling/PerformanceMonitor.Darling.Storage/PayloadDimensions.cs
+++ b/Darling/PerformanceMonitor.Darling.Storage/PayloadDimensions.cs
@@ -111,6 +111,22 @@ public sealed record PayloadDimension(
/// The distinct dimension tables, in creation order.
public static readonly IReadOnlyList DimTables = new[] { QueryTextDimTable, QueryPlanDimTable };
+ ///
+ /// Per dim-feeding fact table, the SQL predicate "this row carries any payload digest" — the OR of
+ /// that table's digest columns from , in declaration order. The #1795 measured GC
+ /// bound probes min(collection_time) under EXACTLY this predicate, and the V39 partial
+ /// indexes are declared with EXACTLY this predicate, so the probe stays an index-edge read instead
+ /// of an oldest-chunk walk. Derived from the map rather than restated so a new dimension column
+ /// lands in the probe the day it lands in — and a pin test holds the V39 DDL to
+ /// these strings so the index predicate can never drift from the probe's.
+ ///
+ public static IReadOnlyDictionary DigestPredicateByTable { get; } =
+ All.GroupBy(d => d.TargetTable, StringComparer.Ordinal)
+ .ToDictionary(
+ g => g.Key,
+ g => string.Join(" OR ", g.Select(d => d.DigestColumn + " IS NOT NULL").Distinct()),
+ StringComparer.Ordinal);
+
/// The diverted columns for one fact table (empty for every table that has none).
public static IReadOnlyList ForTable(string targetTable)
=> All.Where(d => string.Equals(d.TargetTable, targetTable, StringComparison.Ordinal)).ToArray();
diff --git a/Darling/PerformanceMonitor.Darling.Storage/PgMigrations.cs b/Darling/PerformanceMonitor.Darling.Storage/PgMigrations.cs
index 7569acaae6..50cf3d8247 100644
--- a/Darling/PerformanceMonitor.Darling.Storage/PgMigrations.cs
+++ b/Darling/PerformanceMonitor.Darling.Storage/PgMigrations.cs
@@ -82,6 +82,7 @@ public Migration(int version, string name, string sql)
new Migration(36, "ag-latency-columns", V36Sql),
new Migration(37, "ag-local-replica-and-disconnect-refire", V37Sql),
new Migration(38, "query-payload-dimensions", PgSchemaGenerator.GenerateV38PayloadDimensions()),
+ new Migration(39, "dim-feeding-fact-floor-indexes", V39Sql),
};
///
@@ -614,6 +615,29 @@ ALTER TABLE collect.ag_replica_states
ALTER TABLE config.config_alert_settings
ADD COLUMN IF NOT EXISTS ag_disconnect_refire_minutes integer NOT NULL DEFAULT 0;";
+ ///
+ /// V39 — the partial indexes behind the dimension GC's MEASURED bound (#1795). The GC prunes
+ /// dimension content against the oldest SURVIVING digest-carrying fact row rather than an assumed
+ /// horizon, which requires a min(collection_time) probe per dim-feeding table every sweep.
+ /// Unindexed, that is an ordered walk from each hypertable's oldest chunk filtering out NULL-digest
+ /// rows (every pre-#1767 row); with a partial index whose predicate EXACTLY matches the probe's
+ /// WHERE clause, it is an index-edge read. One index per dim-feeding fact table, predicate =
+ /// "carries any digest" (the OR of that table's digest columns from PayloadDimensions.All —
+ /// pinned against the map by test, since a new dimension column would silently fall out of both the
+ /// index and the probe together or neither). TimescaleDB propagates the index to existing and
+ /// future chunks; IF NOT EXISTS keeps the migration idempotent; plain-PostgreSQL stores take
+ /// the same DDL unchanged. Bare names resolve through the migrate session's
+ /// search_path = collect, config, public.
+ ///
+ private const string V39Sql = @"
+CREATE INDEX IF NOT EXISTS ix_query_stats_digest_floor
+ ON query_stats (collection_time)
+ WHERE query_text_digest IS NOT NULL OR query_plan_digest IS NOT NULL;
+
+CREATE INDEX IF NOT EXISTS ix_procedure_stats_digest_floor
+ ON procedure_stats (collection_time)
+ WHERE query_plan_digest IS NOT NULL;";
+
///
/// V9 — the FinOps copy-parity fields that were user-input config or previously live-only:
/// server_properties gains the three inventory columns the shared ServerPropertiesCollector now
diff --git a/Darling/PerformanceMonitor.Darling.Storage/StorageVersion.cs b/Darling/PerformanceMonitor.Darling.Storage/StorageVersion.cs
index 513ae74dc3..31515762a9 100644
--- a/Darling/PerformanceMonitor.Darling.Storage/StorageVersion.cs
+++ b/Darling/PerformanceMonitor.Darling.Storage/StorageVersion.cs
@@ -16,5 +16,5 @@ namespace PerformanceMonitor.Darling.Storage;
///
public static class StorageVersion
{
- public const int SchemaVersion = 38;
+ public const int SchemaVersion = 39;
}
diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.cs
index 5edcf1a3a8..46c911765c 100644
--- a/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.cs
+++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.cs
@@ -399,7 +399,8 @@ OR NOT EXISTS (SELECT 1 FROM pg_extension WHERE extname = 'timescaledb')
EXISTS (SELECT 1 FROM information_schema.columns WHERE table_name = 'config_alert_settings' AND column_name = 'notify_ag_health'),
EXISTS (SELECT 1 FROM information_schema.columns WHERE table_name = 'config_alert_settings' AND column_name = 'ag_disconnect_refire_minutes'),
EXISTS (SELECT 1 FROM information_schema.columns WHERE table_name = 'ag_database_replica_states' AND column_name = 'est_send_drain_time_min'),
- EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'query_plan_dim')";
+ EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'query_plan_dim'),
+ EXISTS (SELECT 1 FROM pg_indexes WHERE indexname = 'ix_query_stats_digest_floor')";
/// The store schema version this viewer build requires — the highest migration it knows
/// (). The connect-time gate blocks a store below this.
@@ -420,7 +421,7 @@ OR NOT EXISTS (SELECT 1 FROM pg_extension WHERE extname = 'timescaledb')
await using var reader = await command.ExecuteReaderAsync(cancellationToken);
if (await reader.ReadAsync(cancellationToken))
{
- return MapProbedSchemaVersion(reader.GetBoolean(0), reader.GetBoolean(1), reader.GetBoolean(2), reader.GetBoolean(3), reader.GetBoolean(4), reader.GetBoolean(5), reader.GetBoolean(6), reader.GetBoolean(7), reader.GetBoolean(8), reader.GetBoolean(9), reader.GetBoolean(10), reader.GetBoolean(11), reader.GetBoolean(12), reader.GetBoolean(13), reader.GetBoolean(14), reader.GetBoolean(15), reader.GetBoolean(16), reader.GetBoolean(17), reader.GetBoolean(18), reader.GetBoolean(19), reader.GetBoolean(20), reader.GetBoolean(21));
+ return MapProbedSchemaVersion(reader.GetBoolean(0), reader.GetBoolean(1), reader.GetBoolean(2), reader.GetBoolean(3), reader.GetBoolean(4), reader.GetBoolean(5), reader.GetBoolean(6), reader.GetBoolean(7), reader.GetBoolean(8), reader.GetBoolean(9), reader.GetBoolean(10), reader.GetBoolean(11), reader.GetBoolean(12), reader.GetBoolean(13), reader.GetBoolean(14), reader.GetBoolean(15), reader.GetBoolean(16), reader.GetBoolean(17), reader.GetBoolean(18), reader.GetBoolean(19), reader.GetBoolean(20), reader.GetBoolean(21), reader.GetBoolean(22));
}
return null;
@@ -445,8 +446,18 @@ OR NOT EXISTS (SELECT 1 FROM pg_extension WHERE extname = 'timescaledb')
/// is unit-tested without a live store; any schema bump past the newest arm trips the pinning test that keeps
/// this in step with .
///
- internal static int MapProbedSchemaVersion(bool hasConfigControlPlane, bool hasAlertDeliveryOverride, bool hasAnalysisState, bool hasAlertTuningKnobs, bool hasDefaultTraceEvents, bool hasIndexObjectStatsLatestIndex, bool hasCollectionLogHypertableOrPlainPg, bool hasJobHistory, bool hasAgentStatus, bool hasGenericWebhook, bool hasDeadlocksDatabaseName, bool hasQueryStoreReplicaRole, bool hasLongQueryCompletions, bool hasWebDashboardConfig, bool hasCustomViews, bool hasServerTags, bool hasConnectionRefireKnobs = false, bool hasAgCollectors = false, bool hasAgAlertKnobs = false, bool hasAgLatencyColumns = false, bool hasAgDisconnectRefire = false, bool hasPayloadDimensions = false)
+ internal static int MapProbedSchemaVersion(bool hasConfigControlPlane, bool hasAlertDeliveryOverride, bool hasAnalysisState, bool hasAlertTuningKnobs, bool hasDefaultTraceEvents, bool hasIndexObjectStatsLatestIndex, bool hasCollectionLogHypertableOrPlainPg, bool hasJobHistory, bool hasAgentStatus, bool hasGenericWebhook, bool hasDeadlocksDatabaseName, bool hasQueryStoreReplicaRole, bool hasLongQueryCompletions, bool hasWebDashboardConfig, bool hasCustomViews, bool hasServerTags, bool hasConnectionRefireKnobs = false, bool hasAgCollectors = false, bool hasAgAlertKnobs = false, bool hasAgLatencyColumns = false, bool hasAgDisconnectRefire = false, bool hasPayloadDimensions = false, bool hasDimFloorIndexes = false)
{
+ /* V39 (#1795 dimension GC measured bound): index-existence sentinel, newest-first arm — the
+ same pg_indexes idiom as the V22 sentinel. ix_query_stats_digest_floor exists only at V39
+ or later. The viewer itself never reads the index; the arm exists so a fully-migrated V39
+ store maps to exactly RequiredStoreSchemaVersion instead of capping at 38 and tripping the
+ connect-time gate against a perfectly healthy store. */
+ if (hasDimFloorIndexes)
+ {
+ return 39;
+ }
+
/* V38 (#1767 query payload dimensions): table-existence sentinel, newest-first arm.
query_plan_dim exists only at V38 or later. The viewer MUST gate on it: at V38 the
collectors stop writing query_text/query_plan_xml inline, so a pre-V38 viewer pointed at a