Skip to content

Legacy trio off the refresh grid (#3653) - #4186

Merged
erikdarlingdata merged 29 commits into
devfrom
fix/3653-a6-freeze
Sep 26, 2026
Merged

erikdarlingdata merged 29 commits into
devfrom
fix/3653-a6-freeze

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

This PR freezes six legacy rollup views so they stop refreshing: query_stats_hourly, procedure_stats_hourly, query_stats_db_hourly, and their three hierarchical dailies. Their interval-honest successors (Q12 hourlies, A6 LB dailies) are already shipped and already cover the same history. The freeze moves purge and retention coverage onto those successors. It also fixes a read-routing regression the freeze itself caused. It closes a gap in the raw-purge gate too, at the seam between the frozen legacy data and the successor's own history.

Built in several passes on fix/3653-a6-freeze. Two review passes, a design review and a security and data-loss review, ran at commit 7d181c99. Every High and Medium finding they raised is fixed on the branch. A third pass then reviewed those fixes at commit 35a0aefb. Its High, a way for the purge gate to release early, is fixed in commit c3291f66. Its other findings are written up in the code as accepted limits or tracked in follow-up issues (see Review).

Part of #3653

What changed

  • The freeze (commit 02a1f124). Six legacy rollups move out of HourlyAggregates and DailyAggregates. They go into a new list, FrozenRollupAggregates: query_stats_hourly, procedure_stats_hourly, query_stats_db_hourly, and their three hierarchical dailies. EnsureContinuousAggregatesAsync still creates all six, on a fresh store or one upgrading past this PR. But it detaches any refresh policy an existing store already has on them, every start, and it never attaches a new one. RollupBackfill.Targets, and the materialization-hole repair targets built from it, exclude the six. Nothing ever backfills or repairs a rollup that no longer advances. Emptying that list broke a separate lookup in DailySummarySql, which then threw for any read naming a frozen view. Commit 8b632d0b added its own lookup, RollupCoverageProbeTargets, which still includes all six, and fixed it.

  • Purge and retention coverage moved to the successors (commit ed49daf6). RawTierCoverage's query_stats and procedure_stats rows named the frozen legacy hourlies as their raw-purge coverage. They now name the interval-honest successors instead, through a new RequireSuccessorOf lookup. query_stats's row also names query_stats_db_hourly's own successor, added in commit 8b632d0b. query_stats_db_interval_hourly reads raw query_stats directly too, so it needed the same coverage. RetentionPolicies' three successor-hourly rows named the legacy daily as their own retention coverage. They now name their own successor daily instead, through SuccessorDailyOf.

  • The stitched floor: purge gate and read routing (commits 3cfeeea2, 0174f585). A successor's own history starts empty on an upgrading store. That broke two things independently. The raw-purge gate held forever, with no self-exit. Every FinOps, health and MCP reader that resolves a tier through RollupCoverage.For() fell back to raw, for any window older than about 4 days. Both are fixed the same way. Each stitches the legacy floor onto the successor's own floor, and takes whichever floor is deeper (earlier). The gate does this in RetentionArmSafetySql. Reads do it in RollupCoverage.For().

  • The seam probe and the hole-walk repair (commit 3098a9e5). The stitch above was not gap-free. An earlier version of its own doc comment claimed it was, and that claim was false. A store can stay stopped for more than a day across the upgrade. It then leaves a raw tail between the legacy's last bucket and the successor's floor. The tail is as long as the stop, minus the one day that the successor's first refresh reaches back. Neither side materializes that tail, and without a repair the raw purge deletes it. RetentionArmSafetySql now probes raw itself, for a row in the tail, before it trusts the stitched floor. RepairMaterializationHolesAsync's scan now reaches back to the legacy's last bucket too, not just the successor's own floor. The tail then reads as an ordinary hole. The service repairs it at startup, up to 24 hourly buckets per start. After an outage of more than a day across the upgrade, the repair starts at the second start (#4186 gate follow-ups: release needs a second start after an outage, plus three smaller holds #4300). Commit c3291f66 makes the repair walk the tail from its newest bucket down, and stop at the first range that fails. Before that, the walk took the oldest buckets first. A partial repair then moved the successor's floor below tail rows that were still missing, and the gate read Covered too early. Commit c14b46a0 gives the tail its own scan window. Before it, the scan stopped at a 4-day horizon, so a tail older than that was never repaired, and the gate held for good.

  • The zero-interval source filter (commit 3cfeeea2). Raw's very first collection row, on any store, always has sample_interval_seconds = 0. The successors already exclude that row from their own history. The gate's own source_oldest check now excludes it too. That one row can no longer hold the gate open past where the successors themselves start.

  • The compression band (commits 4beeb261, 0cb21885). CompressionDeferredUntilFreeze is gone. That was the set that held the 3 successor dailies out of the compression band, until this freeze shipped. AggregateCompressionTargets is 20 members now: 6 hourly, 7 daily, 7 baseline, nothing excluded. (dev has 23 today. The legacy trio still counts there, and the 3 successor dailies are still deferred.) The legacy trio also left HourlyAggregates and DailyAggregates, not just the compression list. So every one of the 17 compression policies that exist today moves to a new band hour on the first start after this ships. Each is counted from AggregateCompressionBandFirstHour. No aggregate's compress_after window changes. Only its hour does. The 3 successor dailies, query_stats_interval_daily, procedure_stats_interval_daily and query_stats_db_interval_daily, are newly compression-registered, at hours 11 to 13. FrozenDailyCompressionDrainStateSql now scopes to all 6 frozen views, not just the 3 dailies. DrainFrozenDailyCompressionPoliciesAsync then drains a pre-freeze store's frozen HOURLY compression jobs too, once every chunk on that view is compressed.

  • The Medium review fix (commit 0cb21885). ConvergeCompressionScheduleAsync runs every start. It only skips views still listed in AggregateCompressionTargets. The frozen six left that list. So an existing store's once-a-day compression jobs on them were being reset to the raw tier's 1-hour cadence, on every start. The converge now skips IsFrozenRollupAggregate views first, unconditionally, before that check. The same commit also widened the drain's view filter to all 6 frozen views (above). It fixed a debug log too, which was still counting the frozen views inside the raw-table bucket.

Review

Two review passes ran at commit 7d181c99: a design review and a security and data-loss review. Every fix below came after both passes. Each fix was checked against the code; no separate reviewer has read them.

High:

  • The raw purge for query_stats and procedure_stats held on every field upgrade (design review). A successor's history starts 1 day before the successor exists, raw keeps 4 days, and only --backfill-rollups released the gate. The fix is a stitched floor with no gap. Commit 3cfeeea2 stitches the legacy floor onto the successor's floor in RetentionArmSafetySql. Commit 3098a9e5 adds the seam probe and the hole-walk repair. Commit c14b46a0 fixes a gap in 3098a9e5 found during the fix check: its repair scan was still limited to the 4-day horizon.
  • The gate compared raw's oldest row against successors that leave out rows with sample_interval_seconds = 0, and the collector writes 0 on its first pass (security review). Commit 3cfeeea2 applies the same filter to the gate's source_oldest check.
  • Every reader that picks a tier through RollupCoverage.For() read the legacy floor alone (security review). Windows older than about 4 days fell back to raw. Commit 0174f585 stitches the floors there the same way the gate does.

Medium:

  • ConvergeCompressionScheduleAsync runs every start and skips only views in AggregateCompressionTargets. The frozen six left that list, so their once-a-day compression jobs were reset to the raw tier's 1-hour cadence on every start (design review). Commit 0cb21885 skips frozen views first. The same commit widens the drain to all 6 frozen views and fixes a debug log that counted them as raw tables.

Low:

  • A doc comment said four frozen views keep a drop_chunks policy. It is three. Fixed in commit 22cbb028.
  • The warning for a failed policy detach on a frozen view said reads fall back to raw scans. In fact the view keeps refreshing on its old schedule until a later start detaches it. Fixed in commit 22cbb028. Commit 35a0aefb then made the warning name a failed create too, because the same catch sees one.
  • Nothing drained the three frozen hourlies' compression jobs. Commit 0cb21885's wider drain covers them now.
  • Successor-hourly retention re-holds, with an alert, only when the successor hourlies predate their dailies by more than 3 days without a backfill. A backfill always releases it. No change needed.
  • Only a manual decompress_chunk run between the drain's check and its removal can leave one chunk uncompressed. That costs storage, not data. No change needed.
  • The refresh-policy detach can run again safely, and a failure affects only its own view. No product path adds a detached policy back. No change needed.

Round 3, at 35a0aefb

A third pass read the fix commits and found 2 High, 2 Medium and 3 Low findings.

Known limits

  • The public release (main, 2026-09-17) has no hole walk. A store that upgrades straight from it keeps any hole inside the legacy rollups from an outage in its last 4 days. After the freeze nothing repairs a legacy rollup, and the raw purge then deletes those rows. A store on a dev build with the hole walk repaired such holes at each start, up to 24 buckets per view per start.
  • The stitched gate trusts the frozen legacy for its whole span, from its oldest bucket to its newest. It checks for a gap only after the legacy's newest bucket. This is round 3's H2, accepted as 3.8.0 parity, and Raw purge gate: check the frozen legacy's span bucket by bucket, and repair holes inside it #4301 tracks the fix.
  • PostgreSQL can run an armed raw purge at its own start and drop a seam older than 4 days before the service holds it (An armed raw purge runs at PostgreSQL start, before the service can hold it, and drops an unrepaired seam #4299).
  • The hole walk runs once per service start and repairs up to 24 hourly buckets per view. So a seam several days long takes that many starts to close. The seam and the rest of the walk share those 24 buckets, and the seam goes first. Until it closes, the raw purge for query_stats and procedure_stats stays held, and raw keeps growing. --backfill-rollups closes it in one run.

Tests

Four live tests in FrozenRollupLiveTests.cs drive the product's own start-up entry points against a real TimescaleDB instance, not a re-implementation of them:

  • DroppingFrozenHourlyChunks_LeavesItsFrozenDailyUnchanged drives EnsureContinuousAggregatesAsync and RepairMaterializationHolesAsync.
  • StartPathRunTwice_NeverReAddsAFrozenRefreshPolicy_AndEveryOtherAggregateKeepsOne drives EnsureContinuousAggregatesAsync, run twice.
  • RawPurge_ArmsOffSuccessorHourlyCoverage_NotTheEmptyLegacyOnes drives EnsureRetentionPoliciesAsync.
  • FrozenDailyCompressionPolicy_DrainsOnlyOnceEveryChunkIsCompressed drives EnsureAggregateCompressionAsync.

Each has a revert-proof. The author reverted the guarding code, watched the test fail with the expected error, then restored the code and watched it pass. The seam fix (commit 3098a9e5) adds two more live tests the same way, Outage_SeamBetweenFrozenLegacyAndSuccessor_HoleWalkRepairsItAndGateReleases and Outage_EmptySeam_GateReportsCoveredWithNoHoleWalk, plus a pure pin, RetentionArmSafetySql_StitchedSlot_ProbesSeamBeforeFallingBackToLegacyFloor. All three carry the same revert-then-restore proof. A fourth live test from the High fix, FieldUpgrade_EmptySuccessors_LegacyFilled_ReportsRawPurgeCovered, needs DARLING_TEST_PG to run and has no revert-proof recorded in the comments.

Commit c3291f66 adds a pure pin, TheCap_NewestFirst_TakesTheNewestFirst_SplitsAStraddlingRangeAtItsOlderEdge_AndDefersTheRest, and two live tests:

  • Outage_SeamWiderThanTheCap_NewestFirstRepairsTheTopAndKeepsTheGateHeldUntilFullyRepaired: a 35-bucket seam keeps the gate held after one walk, and releases it after the second. With the old oldest-first order, it failed at the first gate check.
  • Outage_SeamTwoRanges_NewerRangeFails_OlderRangeUntouchedAndGateStillNotSafe: a CHECK constraint makes the newer range's refresh fail. The older range is never touched, and the gate stays held.

The changed classes were run, all green, but not the full suite. CI runs the full suite.

Commit c14b46a0 adds five pure tests for its new MaterializationHoleScanWindows helper, and a live twin of the seam test for a 6-day outage, Outage_SeamOlderThanTheHorizon_HoleWalkStillRepairsItAndGateReleases. With the old horizon limit put back, the new live test fails ("expected the 2-bucket seam tail to be repaired, got 0"). With the fix restored, it passes. On a fresh local rig, FrozenRollupLiveTests (8 tests), TimescaleContinuousAggregateTests (41) and MaterializationHoleRepairTests (16) passed with no failures, before and after a merge of dev. The full suite did not run locally after this commit. GitHub Actions CI passed on 35a0aefb and again on the merged head fd53564c: build, Darling PostgreSQL tests and Lite tests.

Every test the freeze broke is fixed. This body drops the original test-status inventory, because none of it is current any more.

The last full local run of Darling.Tests was on 3098a9e5, before the horizon fix. It reported 13760 total tests: 0 errors, 2 failed, 47 skipped, 1 not run, in 553 seconds. Both failures, ServerListAndSummaryPlanShapeTests and CaptureDownChunkOrderTests, are chunk-visitation-count plan-shape checks in a subsystem this PR does not touch.

Whether both are pre-existing was not confirmed locally. The CI run against the merged dev (see the merge gate) is the check.

Merge gate

The earlier prerequisites are merged: LA (#4182), LB (#4181) and LA-8 (#4184).

Item Status
i. The install and the A6 successor-daily backfill on every store (runbook docs/runbooks/a6-successor-daily-backfill.md). Done.
ii. Floor check: min(bucket) of query_stats_interval_hourly and of query_stats_db_interval_hourly is at or before raw query_stats' oldest row. Passed per store after backfill (all SQL Server stores; not applicable to the PostgreSQL store).
iii. Live stitched reads show no gap and no overlap. Stitched reads verified per store after backfill: no gap, no overlap; daily step 4 excludes the refresh lag.
iv. The review round. Two passes ran at 7d181c99, and every High and Medium they found is fixed. A third pass ran at 35a0aefb. Its H1 is fixed in c3291f66, and that fix was checked against the code. Its H2 is accepted as 3.8.0 parity, and #4299, #4300 and #4301 track the rest. No fourth round runs.

CHANGELOG entry

SECTION: Changed
ENTRY:

  • Three legacy query-stats rollups stop refreshing ([Legacy trio off the refresh grid (#3653) #4186]) - query_stats_hourly, procedure_stats_hourly, query_stats_db_hourly and their daily rollups keep the history they already hold. Their interval-honest successors take over. Reads older than the successors' own history still use them. New data goes only to the successors, so the store no longer spends refresh time on the old rollups. The raw query_stats and procedure_stats purge now waits until the successors and the old rollups together cover every raw row. After a long outage, the service fills the gap at startup, up to a day of it per start. Only then does the purge resume. --backfill-rollups fills it in one run. Existing compression jobs move to new hours once. Dashboards, alerts and queries do not change.
    REF:
    [Legacy trio off the refresh grid (#3653) #4186]: Legacy trio off the refresh grid (#3653) #4186

This entry goes into CHANGELOG.md in the batched splice after this PR merges.

erikdarlingdata and others added 4 commits September 24, 2026 21:33
… off the refresh grid

Moves query_stats_hourly, procedure_stats_hourly, query_stats_db_hourly and their three
legacy dailies out of HourlyAggregates/DailyAggregates into a new FrozenRollupAggregates
list (OffGridAggregates pattern, but with no policy builder at all): the ensure sweep still
creates a missing one and now actively detaches any refresh policy an existing store still
carries on it, and no converge can re-add one since none of the policy-builder lists name it
anymore. RollupBackfill.Targets excludes the frozen six so MaterializationHoleTargets' Single()
lookup against Hourly/DailyAggregates cannot throw for a view neither list has anymore.

Design points 3, 4, 6 and 7 (raw-purge coverage, successor-hourly retention coverage, the
CompressionDeferredUntilFreeze removal, and the frozen-daily compression drain rule) are not
yet done — see the PR body for the full handoff.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
… band, drain the frozen dailies' policies

Point 6: deletes CompressionDeferredUntilFreeze and its .Where filter, so
AggregateCompressionTargets is every member of HourlyAggregates, DailyAggregates and
BaselineAggregates with nothing subtracted -- 20 (6 + 7 + 7), not 17. The three
interval-honest successor dailies move from no compression policy to hours 11-13 on the
band; the seven baseline aggregates shift from hours 11-17 to 14-20. The six hourly
members and the four pre-existing daily members keep their hours.

Point 7, the drain: EnsureAggregateCompressionAsync's converge already left every
non-target aggregate's policy alone (confirmed by reading it -- non-targets are filtered
out of the state read entirely, so nothing downstream ever touches them). Adds
DrainFrozenDailyCompressionPoliciesAsync, called as a last step in
EnsureAggregateCompressionAsync: a probe scoped by name to the three frozen legacy
dailies that counts every uncompressed chunk on their materializations (not the
age-gated eligible_under_*_rule columns, which answer a different question), a pure
predicate (a job exists and no chunk is uncompressed means drain), and a
remove_compression_policy call per drained daily, logged. The frozen hourlies get no
drain; their chunks age out through retention.

Re-pins the TimescaleAggregateCompressionTests pins this touches, by derivation rather
than new literals, and adds pure tests for the drain predicate's truth table and the
probe SQL's naming. 17 of the 22 red tests listed in the brief remain red -- refresh-grid
and phase-slot pins in TimescaleSupportTests/RefreshCeilingProvenancePinTests/
RefreshCeilingStalenessTests unrelated to this lane's points, and coverage/raw-gate/
target-count pins in IntervalHonestHourlyRollupTests/MaterializationHoleRepairTests --
left for a follow-up lane per the PR body's handoff.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…at the successors, not the frozen legacy rollups

RawTierCoverage's query_stats/procedure_stats rows and RetentionPolicies'
three successor-hourly rows still named the six rollups LC froze off the
refresh grid. A gate keyed on a frozen relation either finds it empty once
its own retention trims it with nothing refilling it (raw purge, blocks
forever) or never actually asks whether the real consumer caught up
(successor retention, since the legacy daily's floor is fixed and always
looks "covered"). Both are now derived from SupersededHourlyRollups and
SupersededDailyRollups through two new lookups (SuccessorOf's daily
counterpart, and throwing wrappers for callers that already know the
relation is superseded), so the two registries and the two gates cannot
drift apart again. Re-pinned the one IntervalHonestHourlyRollupTests test
that exercised this by name; left the compression-band, refresh-grid,
slot and repair-target pins this touches for their respective owners.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
@erikdarlingdata

Copy link
Copy Markdown
Owner Author

LC-a2 (points 3 and 4: purge and retention coverage)

Commit ed49daf (merged with LC-a3's compression work at 1cc0bb2), pushed to fix/3653-a6-freeze.

What changed

Darling/PerformanceMonitor.Darling.Storage/TimescaleSupport.cs:

  • Point 3, RawTierCoverage (~:6026). The query_stats and procedure_stats rows named the frozen legacy hourlies (QueryStatsHourlyView, ProcedureStatsHourlyView) as their coverage. Both now read RequireSuccessorOf(...), a new wrapper around the existing SuccessorOf lookup. It throws instead of returning null, since these two calls already know the relation is superseded. query_store_stats is untouched: it is not one of the frozen six.
  • Point 4, RetentionPolicies (~:6070). The three successor-hourly rows (QueryStatsIntervalHourlyView, ProcedureStatsIntervalHourlyView, QueryStatsDbIntervalHourlyView, including the database-grain one) named the LEGACY daily as coverage. They now read a new SuccessorDailyOf/RequireSuccessorDailyOf pair that looks up SupersededDailyRollups by SuccessorHourly, so each names its own successor daily.
  • Rewrote the block comment ahead of those three rows. The old text said the successors had "no daily of their own yet," which A6 LB already made false. Rewrote doc summaries elsewhere that still described the old shape. SupersededHourlyRollups' and SupersededDailyRollups' own doc comments already described this end state correctly, written forward-looking by an earlier lane. They needed no changes.
  • Everything else the freeze already did is unmodified: legacy hourlies keep their own retention policy, FrozenRollupAggregates, backfill/repair-target exclusion. That is points 1/2's work, not touched here.

Darling/Darling.Tests/IntervalHonestHourlyRollupTests.cs: re-pinned EverySuccessor_IsInTheCoverageProbe_TheBackfillPlan_AndTheRetentionLadder_ButNotTheRawGate, renamed to ...TheRetentionLadder_AndTheRawGateWhereItHasOne. It asserted the OLD shape by name: retention coverage equal to the legacy daily, and unconditional absence from the raw gate. Points 3 and 4 reverse both of those for two of the three pairs. Its backfill-order check also pointed at the legacy daily, which already left RollupBackfill.Targets under points 1/2, so it threw before reaching the coverage assertion at all. Rewrote it to derive every expected name from the registries (SuccessorOf, SuccessorDailyOf) rather than typing one, and to assert the raw-gate split correctly. The query_stats/procedure_stats successors are now IN RawTierCoverage; query_stats_db's successor stays out, since it has no raw table of its own: it shares query_stats' raw table with query_stats_hourly, already that row's consumer.

Red proof

Direct inspection of commit 02a1f12 (git show 02a1f124:...) shows RawTierCoverage naming QueryStatsHourlyView/ProcedureStatsHourlyView literally, and RetentionPolicies' three successor rows naming QueryStatsDailyView/ProcedureStatsDailyView/QueryStatsDbDailyView literally. Running the 7 named classes against 02a1f12, before any edit, gave 22 failed of 163 total, matching the brief. After my fix, before merging LC-a3: 21 failed of 163 total. Exactly one test moved from fail to pass, and nothing else moved either direction.

Pins fixed vs. left (after merging LC-a3, current HEAD, 165 total)

Fixed (mine): 1, IntervalHonestHourlyRollupTests.EverySuccessor_IsInTheCoverageProbe_TheBackfillPlan_TheRetentionLadder_AndTheRawGateWhereItHasOne.

Left, by class and reason:

  • TimescaleSupportTests: 7 failed of 69 total, 16 skipped. All refresh-grid/slot (CompressionPhaseGrid..., TheRejectedWatchLineAlternative..., TheRefreshSlotWatchLines..., NoCompressionMinuteStarts..., TheJobCadenceKnob..., TheRefreshGridIsUnchanged..., TheRefreshSlotReading...). Left for LC-a3 per brief.
  • RefreshCeilingProvenancePinTests: 3 failed of 16 total. Refresh-grid/slot. Left for LC-a3.
  • RefreshCeilingStalenessTests: 1 failed of 18 total. Refresh-grid/slot. Left for LC-a3.
  • TimescaleAggregateCompressionTests: 0 failed of 12 total, 2 skipped. LC-a3 already fixed these (5 failed before the merge).
  • MaterializationHoleRepairTests: 2 failed of 10 total. Not coverage/raw-gate. Both reference MaterializationHoleTargets.Single(... == QueryStatsDailyView), the legacy daily that already left that list under points 1/2. One is a stale hardcoded count (26, now 20) plus a stale backfill-order check against the legacy daily. The other calls the repair-scan SQL builder for the legacy daily directly.
  • IntervalHonestHourlyRollupTests: 2 failed of 8 total. The class's 3rd failure, ThreePairs_LegacyStaysRegistered..., is a pure HourlyAggregates count pin (9 vs 6), refresh-grid, left for LC-a3. EachSuccessor_IsItsLegacysShape... and EveryHourlyTierReader_TakesTheSuccessorWhereItCovers... both fail via MaterializationHoleTargets/DailySummarySql referencing the legacy daily. The second is reader code my brief told me not to touch.

Of the 16 remaining, 12 are refresh-grid/slot (LC-a3's, or already fixed by LC-a3), and 4 are repair-target membership debt from points 1/2 that is not explicitly named in anyone's brief here. MaterializationHoleTargets in production already excludes the frozen six correctly, confirmed by the exception text itself. But these 4 tests, and TimescaleContinuousAggregateTests (3 more failures outside my required list, same refresh-grid cause, not touched), were never re-pinned for either the 9-to-6 hourly count or the repair-target exclusion. Flagging this for the coordinator to route. It is pre-existing, confirmed by reading source and by the exception text, not something this change introduced.

Build and tests

Darling.Tests builds with 0 errors and 6 pre-existing warnings, all in SqlServerStoreTileBehaviour*Tests.cs, a file this lane did not touch. Ran every class in both touched files: IntervalHonestHourlyRollupTests, and IntervalHonestHourlyRollupLiveTests (0 failed, 1 skipped, needs DARLING_TEST_PG, not set up here). Also ran all 7 brief-named classes, individually and combined. Spot-checked TimescaleContinuousAggregateTests, PayloadDimensionTests, PgTargetAnomalyTests for regressions: none found. The 3 TimescaleContinuousAggregateTests failures are the same pre-existing refresh-grid cause described above. Did not run the full suite; that is left to LC-b per the brief.

🤖 Generated with Claude Code

https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3

erikdarlingdata and others added 3 commits September 24, 2026 22:08
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…eze's 12-light, 21-minute window (#4186)

The freeze moved the legacy trio off the hourly refresh grid and into
FrozenRollupAggregates, with their successors taking the three positions
the legacy trio held instead of staying appended behind it. That returns
the grid to its pre-Q12 shape (twelve light members, 21-minute heaviest
window, 1,050 s watch line) rather than Q12's fifteen/18/900. Ten pins
across four files pinned the old Q12 numbers; this re-derives each one
against the product's own functions and updates the doc comments to
narrate the append-then-freeze round trip.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
@erikdarlingdata

Copy link
Copy Markdown
Owner Author

LC-a5 (refresh grid re-derived after the freeze)

Ran out of budget partway through the 15-pin list plus the coordinator's two added pure pins. Ten are fixed, green, and pushed (commit 6419e4b7 on fix/3653-a6-freeze). Seven are still red and listed below with the numbers I already worked out, so the next lane does not have to re-derive them.

The freeze returns the hourly grid to its pre-Q12 shape: HourlyAggregates is 6 members again (legacy trio moved to FrozenRollupAggregates, successors take its 3 positions in place rather than staying appended), so HourlyRefreshPhaseOrder is 13 total / 12 light again. LightHourlyRefreshCount=12, LightBandSpanMinutes=11, HeaviestRefreshStartMinute=15, HeaviestRefreshWindowMinutes=21, RefreshPhaseSlotSeconds=1260, RefreshSlotWarningSeconds=1050. Checked by writing each new expected value from that derivation (never from the failing test's own output), then building and running.

Fixed and green (10)

Test Invariant Old (Q12) New (LC) How checked
TimescaleSupportTests.CompressionPhaseGrid_ClearsEveryRefreshSlotsGuardBand_AndTheHeaviestRefreshsSlotWhole light count / band span / heaviest start & window / refresh-minute set / guard & window exclusion counts 15 / 14 / 18 / 18 / {0-14,18} / 18,18 12 / 11 / 15 / 21 / {0-11,15} / 15,21 Ran against live TimescaleSupport properties
TimescaleSupportTests.NoCompressionMinuteStartsWhileTheHeaviestRefreshIsStillRunning ceiling's % and seconds clear of the watch line 0%, 4s 17%, 154s (1050-896)*100/896, 1050-896
TimescaleSupportTests.TheRefreshSlotWatchLines_AreDerivedFromTheWindow_NotWrittenDown slot seconds, watch line, lead time 1080, 900, 180 1260, 1050, 210 Direct read of RefreshPhaseSlotSeconds/RefreshSlotWarningSeconds
TimescaleSupportTests.TheRejectedWatchLineAlternative_... (renamed ..._NoLongerSitsBelowTheRecordedCeiling_SoTheRejectionIsCouplingAgain) alternative (slot - 1 guard band) vs ceiling alternative 840, BELOW ceiling alternative 1020, ABOVE ceiling (896) — ordering argument lost again, coupling is the sole reason, same as pre-Q12 #3174 Recomputed RefreshPhaseSlotSeconds - CompressionPhaseGuardMinutes*60; renamed since old name asserted a now-false "sits below"
TimescaleSupportTests.TheRefreshSlotReading_CarriesItsOwnVerdict_AndReportsOverrunAsNegativeHeadroom headroom at ceiling / at watch line 184s/83.0%, 180s/83.3% 364s/71.1%, 210s/83.3% HeaviestRefreshSlotReading live computation
TimescaleSupportTests.TheJobCadenceKnob_FiresInsideTheWindow_WhichIsWhyTheSlotWatchIsSeparate knob (900s) vs watch line EQUAL (900==900, Q12 coincidence) knob 150s BEFORE watch line (900<1050) — reopened, same as pre-Q12 Replaced equality assert with the reopened inequality + gap
TimescaleSupportTests.TheRefreshGridIsUnchanged_AndTheCompressionGridsOneInputFromItIsPinned full (view, minute) table, 13 rows 16-row table, heaviest :18 13-row table, heaviest :15 (table in PR diff) Hand-derived from RefreshPhaseMinutesFor's own algorithm (unbounded-every-4th, bounded fills gaps), confirmed by running
CompressionPhaseAssignmentTests.TheGridDoesNotMove_TheAssignmentIsAPermutationOfTheSameMinutes HeaviestRefreshWindowMinutes literal 18 21 Direct
BaselineSupplyTests.Successors_TookTheLegacyPositions_SoThePhaseGridDidNotMove registry count, this pair's minutes, heaviest start, watch line 16 / 6,7 / 18 / 900 13 / 3,5 / 15 / 1050 Direct
IntervalHonestHourlyRollupTests.ThreePairs_... (renamed ..._LegacyFrozenOffTheGrid_SuccessorTakesItsHourlyPosition_DailyFreezesBesideTheLegacy) legacy trio's registry membership in HourlyAggregates/DailyAggregates (9/10 members) in FrozenRollupAggregates (6/6), off both grid lists (6/7 members) Rewrote against FrozenRollupAggregates, IsFrozenRollupAggregate

Jobs that re-phase once on the first start of this build (from Q12's minute to LC's, all thirteen survivors move because the legacy trio's removal reflows every position): query_store_stats_hourly :04→:00, query_store_stats_interval_hourly (heaviest) :18→:15, query_store_stats_corrected_hourly :08→:04, query_stats_interval_hourly :12→:08, procedure_stats_interval_hourly :03→:01, query_stats_db_interval_hourly :05→:02, perfmon_interval_baseline :06→:03, wait_stats_interval_baseline :07→:05, session_stats_baseline :09→:06, query_stats_baseline :10→:07, blocked_process_baseline :11→:09, deadlock_baseline :13→:10, memory_baseline :14→:11. The three legacy views (query_stats_hourly :00, procedure_stats_hourly :01, query_stats_db_hourly :02 under Q12) leave the grid entirely rather than moving — their refresh policies are dropped on start, not re-phased.

Totals by class: TimescaleSupportTests 7/7 fixed, CompressionPhaseAssignmentTests 1/1, BaselineSupplyTests 1/1, IntervalHonestHourlyRollupTests 1/1 (the other 2 failures in that class, EachSuccessor_... and EveryHourlyTierReader_..., are LC-a4's per the brief and are still red — not touched).

Build: Darling.Tests 0 warnings, 0 errors. Ran TimescaleSupportTests, CompressionPhaseAssignmentTests, BaselineSupplyTests, IntervalHonestHourlyRollupTests together: Total: 104, Failed: 2 (both LC-a4's), Skipped: 16 (live, no rig per brief).

Still red (7) — not reached, with what I worked out

RefreshCeilingProvenancePinTests (3): EveryDerivationClaim_FollowsFromTheConstantsAndThePublishedPopulation, TheReadingsAboveTheStatedPercentile_BoundTheCensusCount, ThePopulationFloor_ReportsAPopulationTooThinToSupportAMaximum. These parse TimescaleSupport.cs's own doc-comment prose by regex and cross-check the numbers against live constants — the fix is in the product's doc comments, not the test file, except where noted.

  • HeaviestHourlyRefreshObservedCeilingSeconds's doc, "margin sentence": At 896 s against a 1080-second slot the margin is 184 seconds → needs ...1260-second slot the margin is 364 seconds.
  • Same doc, "measured live envelope": ...maximum was 896.1 s. That leaves 183.9 s of the slot, 17.0% of it, and sits 3.9 s BELOW... → needs ...leaves 363.9 s of the slot, 28.8% of it, and sits 153.9 s BELOW... (I derived these four figures; not yet written or verified by a run).
  • RefreshSlotWarningSeconds's doc, rejected-alternative paragraph: currently says the alternative (1,020s at this window) sits BELOW HeaviestHourlyRefreshObservedCeilingSeconds — false now (1020 > 896, matches the TimescaleSupportTests fix above). Needs the prose changed to ABOVE with the 1,020 figure, and the regex in RefreshCeilingProvenancePinTests.cs ("the rejected alternative watch line" pin, matches literal "sits BELOW") updated to match "sits ABOVE" — same wording-moved-with-the-pattern precedent Q12 used (see that pin's own comment).
  • OtherHourlyRefreshObservedCeilingSeconds's doc, scope sentence: it covers 12 of the 15 light views the constant now bounds; the 3 registered after the read are unmeasured. Verify() forces the first number to the historical census's own count (12, fixed forever) and the second to live LightHourlyRefreshCount (now 12 too) — so the formula gives 12 of 12; 0 unmeasured. That is not true: the current 12 light members are 9 that the original census actually measured (2 Query Store + 7 baseline) plus the 3 successors, which are still unmeasured (the 3 legacy views the census did measure are the ones that just left the grid). Writing "0 unmeasured" would contradict the very next sentences, which still argue the 3 successors are safe "by shape, not by a reading." The test's own comment names the exit condition as covered == total retiring the sentence — I'd take that literally (delete the sentence + its pin) rather than write a misleading "0 unmeasured" claim, but flagging it rather than doing it blind, since it's a judgment call on a verification framework, not just a number swap.
  • Did not check CompressionMinutesDeclaration's sentences (occupies X of the Y seconds, admit Y minutes that sit INSIDE the refresh, other Y minutes... past the refresh) — these likely shift with the window (18→21) but I never located/read them.

RefreshCeilingStalenessTests (1): TheMeasuredRunThatDemonstratedTheDefect_StillFalsifiesTheRecordedCeiling. Not investigated. Relevant fact found in passing: RefreshSlotWarningSeconds's doc quotes "a 952 s clean run of 2026-09-08 17:00Z" that at Q12's 900s line classified ApproachingSlot; at LC's 1050s line it's back to InsideSlot (952<1050) — this may be exactly what "falsifies" refers to and may need its own reconciliation.

TimescaleContinuousAggregateTests (3): ContinuousAggregatePolicy_IsTheConservativeHourlyShape_Idempotent, HourlyRefreshPhases_AreDistinctForEveryRelationTwoPoliciesContendFor, HourlyRefreshPolicies_AreStaggeredAcrossTheHour_OnAFixedSchedule. Not investigated at all — no line numbers gathered. Likely the same class of grid-count/minute-list literals as the fixed TimescaleSupportTests pins.

None of my fixes touch a live/rig-gated test; the 16 skipped in my run all need DARLING_TEST_PG, none of which I have.

Plain-English checker: not run (past budget; nothing but this comment to check, and it's short).

No product code was changed — only test files. No files owned by LC-a4 (DailySummarySql.cs, TimescaleSupport.MaterializationHoles.cs, the EachSuccessor_.../EveryHourlyTierReader_.../EverySuccessor_IsInTheCoverageProbe_... methods) or LC-b were touched.

erikdarlingdata and others added 7 commits September 24, 2026 22:34
…y rollup, repair-list pins re-derived, raw gate covers query_stats_db_interval_hourly

DailySummarySql.QueriesCteForCagg/QueriesCteForStitchedCagg looked their
relation up in TimescaleSupport.MaterializationHoleTargets, which the freeze
(LC) emptied of the six frozen legacy rollups. Every daily-summary read that
named a frozen view (legacy-only, RollupCoverage.Unknown, or a stitched splice
below the successor floor) threw ArgumentException.

Added TimescaleSupport.RollupCoverageProbeTargets: the same per-relation
descriptor as MaterializationHoleTargets, but over every RollupViews member
including the frozen six, with CREATE text sourced from HourlyAggregates,
DailyAggregates or FrozenRollupAggregates. Pointed the three DailySummarySql
lookups at it. MaterializationHoleTargets itself is untouched, so the repair
walk still never sees a frozen view.

Re-pinned MaterializationHoleRepairTests' two lookups that named
QueryStatsDailyView directly (now frozen) to its live successor via
SuccessorDailyOf, re-derived the stale registered-count literal (26 -> 20)
and the daily-after-hourly ordering check (now over the live successor pair,
not the frozen legacy pair), and added a pin that MaterializationHoleTargets
holds no frozen view. Re-pinned the two MaterializationHoleRepairLiveTests
that built their whole scenario on the now-frozen query_stats_hourly to use
its live successor instead; the one scenario that compared a live unfiltered
rollup against a live filtered one no longer has two live sides after the
freeze, so it now proves only the half that is still live: a restart-only
hour is neither materialized nor reported as a hole.

RawTierCoverage gated the raw query_stats purge on query_stats_interval_hourly
only, but query_stats_db_interval_hourly also reads raw collect.query_stats
directly. Added it through RequireSuccessorOf(QueryStatsDbHourlyView), fixed
the pin that asserted it stayed out, and added a derived pin that walks every
RawTierCoverage row and checks its coverage set against every non-frozen
RollupViews entry sourced from that raw table directly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…ollup by its lookup

MaterializationHoleScanShapeLiveTests and DailySummaryReadShapeLiveTests both
looked a frozen legacy view (query_stats_hourly/_daily) up in
MaterializationHoleTargets directly, the same crash pattern as the daily
summary's own lookup. Re-pinned the scan-shape oracle to the live successor
(SeedAsync already refreshes it over the same windows) and dropped the
daily-tier case, whose only live target is now frozen; re-pinned the
read-shape oracle to RollupCoverageProbeTargets, mirroring the product's own
fix. Compiles clean; not yet run against a live rig in this session.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
Four live TimescaleDB tests against the LC freeze: dropping a frozen
legacy hourly's chunks leaves its frozen daily's rows and sums
unchanged, no start-path converge ever re-adds a frozen refresh
policy while every other aggregate keeps one, the raw purge arms off
the successor hourlies' coverage even though the legacy hourlies are
empty (and stays held when a coverage relation falls short), and a
frozen daily's compression policy drains only once every chunk it
will ever hold is compressed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
The legacy six moved to FrozenRollupAggregates, left RollupBackfill.Targets
and HourlyRefreshPhaseOrder, and stopped getting a refresh policy. Every
test that used one of them as a convenient live example now moves to its
interval-honest successor; every test that counts EnsureContinuousAggregatesAsync's
"ready" total now adds FrozenRollupAggregates.Length; the compression pin
moves from 23 to 20 targets, and the successor dailies now get a real
compression policy instead of the removed deferral.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
@erikdarlingdata

Copy link
Copy Markdown
Owner Author

LC-a4 (the daily summary crash, the lookup audit, the raw gate)

Item 1: the daily summary crash

DailySummarySql.QueriesCteForCagg/QueriesCteForStitchedCagg looked their relation up in
TimescaleSupport.MaterializationHoleTargets, which the freeze emptied of the six frozen legacy rollups. Every
daily-summary read naming a frozen view (legacy-only, RollupCoverage.Unknown, or a stitched splice below the
successor floor) threw ArgumentException.

Fix: added TimescaleSupport.RollupCoverageProbeTargets (TimescaleSupport.MaterializationHoles.cs), the same
per-relation descriptor as MaterializationHoleTargets, built from every RollupViews member including the
frozen six. Its CREATE text comes from HourlyAggregates, DailyAggregates or FrozenRollupAggregates; together
they hold exactly one entry per RollupViews member. Pointed all three DailySummarySql lookups at it and
updated their refusal message. MaterializationHoleTargets itself is untouched, so the repair walk still never
sees a frozen view.

Red proof: IntervalHonestHourlyRollupTests.EachSuccessor_IsItsLegacysShape_... and
EveryHourlyTierReader_TakesTheSuccessorWhereItCovers_... both failed on 1abc48b (Sequence contains no matching element, or the daily summary throwing) and pass after. Added a new pin,
DailySummaryStitchedRangeTests.FrozenLegacy_StillBuilds_UnderUnknownCoverage_AndUnderAStitchThatCrossesTheWindow.
For both frozen names this overload ever names, query_stats_hourly at the Hourly tier and query_stats_daily
at the Daily tier, RangeSqlFor now builds under RollupCoverage.Unknown and under a stitch whose successor
floor falls inside the window. It throws on 1abc48b and passes after.

Coordinator's correction reported 54 CI failures, not 19. Confirmed and fixed on the pure side:
DailySummaryStitchedRangeTests (9/9), DailySummaryNotCarriedTests (two more instances of the same lookup
pattern in local test helpers, both re-pinned to RollupCoverageProbeTargets), DailySummaryReadShapeSqlTests,
RetentionTierRouterTests. All green locally. On the live side, I ran out of context budget before I could stand
up the PG rig (see Deferred below). I did find and fix one more instance by static derivation, not rig-verified
this session: DailySummaryReadShapeLiveTests had the identical MaterializationHoleTargets.Single oracle
lookup, re-pinned to RollupCoverageProbeTargets.

Item 2: the same bug elsewhere

Grepped production .cs under Darling/ (excluding Darling.Tests) for .Single/.First/.FirstOrDefault on
HourlyAggregates, DailyAggregates, RollupBackfill.Targets, MaterializationHoleTargets,
AggregateCompressionTargets, RollupViews. Zero hits outside DailySummarySql.cs (fixed above). Every other
production site iterates these lists rather than looking up one member by name.

Checked RefreshPhaseMinutesFor (TimescaleSupport.cs :3261/:3318) by name: it throws for a view absent from
HourlyRefreshPhaseOrder, which derives from HourlyAggregates union BaselineAggregates only. No production
path calls it with a frozen view's name, because the frozen six get no refresh policy at all. Nothing ever asks
for their phase minute. Verdict: safe, no fix needed.

Item 3: MaterializationHoleRepairTests (2 red)

Both named a frozen QueryStatsDailyView in a .Single lookup. Re-pinned
TheSourceFilter_IsTheAggregatesOwnWhere_Verbatim_OrEmpty's unfiltered-daily case to the live successor via
SuccessorDailyOf (derived, not typed); its CREATE is equally unfiltered. Re-derived
Targets_AreEveryRegisteredAggregate_...'s stale count, 26 to 20, since HourlyAggregates plus DailyAggregates
lost six members, and re-derived its daily-after-hourly ordering check. That check now runs over the live
SupersededDailyRollups pair instead of the frozen SupersededHourlyRollups pair, which is no longer in the
repaired set at all. Added MaterializationHoleTargets_HoldsNoFrozenRollupAggregate, asserting no member
satisfies IsFrozenRollupAggregate.

I also found and fixed the same pattern in the live classes the coordinator's correction named, which are not at
the brief's original line numbers: MaterializationHoleRepairLiveTests.PreOutageTail_... and
OneHoleRepairedAndOneDeferred_... both built their whole scenario on query_stats_hourly directly. Re-pinned
both to the live successor. PreOutageTail_... also had a "control on the source filter" section contrasting a
live unfiltered relation against the live filtered successor. The freeze removed the only unfiltered live
relation in this family, so that contrast no longer has two live sides (the pure CreateSql contrast still holds
in MaterializationHoleRepairTests). I narrowed the control to what remains live and true: a restart-only hour is
neither materialized nor reported as a hole for the successor. MaterializationHoleScanShapeLiveTests had the
same two frozen lookups. Re-pinned TheFencedScan_...'s hourly case to the successor and dropped its daily-tier
case, since no live daily target remains without extending the seed, which I did not do (see Deferred). Re-pinned
TheScanProbesPerBucket_... outright.

None of these four live methods ran against a rig this session. All compile, and all are mechanical,
derivation-based re-pins of what the pure pins already prove.

Item 4: the raw purge gate

RawTierCoverage's "query_stats" row named only RequireSuccessorOf(QueryStatsHourlyView), but
query_stats_db_interval_hourly also reads raw collect.query_stats directly
(CreateQueryStatsDbIntervalHourlySql). Added RequireSuccessorOf(QueryStatsDbHourlyView) to that row.

Added RawTierCoverage_IsExactlyEveryNonFrozenRollupThatReadsThatRawTableDirectly, which derives each row's
expected coverage as every non-frozen RollupViews entry whose Source is that row's raw table, and checks it
against the row. It confirms both the fix and that the existing procedure_stats/query_store_stats rows are
already complete: no second missing consumer turned up.

Fixed EverySuccessor_IsInTheCoverageProbe_TheBackfillPlan_TheRetentionLadder_AndTheRawGateWhereItHasOne, which
had asserted the db-grain successor stayed out of every coverage row. It now asserts the successor is in
query_stats's row and nowhere else. Updated the doc comments on both sides.

Totals

Local run, pure classes, Darling.Tests.exe: IntervalHonestHourlyRollupTests, RollupBackfillTests,
MaterializationHoleRepairTests (pure half), DailySummaryNotCarriedTests, DailySummaryStitchedRangeTests,
DailySummaryReadShapeSqlTests, RetentionTierRouterTests. 101 run, 101 passed after merging LC-a5's own push.
One failure before that merge, ThreePairs_LegacyStaysRegistered_..., was LC-a5's own pin, untouched by me, and
went green once their fix landed. Build: Darling.Tests.csproj, 0 Warning(s), 0 Error(s).

Deferred: context budget, did not reach the PG rig this session

  • The rig never came up, so none of the live classes ran. MaterializationHoleRepairLiveTests,
    MaterializationHoleScanShapeLiveTests, and the one method I fixed in DailySummaryReadShapeLiveTests are
    unverified beyond compiling.
  • Not inspected at all this session: DailyStitchLiveTests, DailySummaryAndComposeStitchedLiveTests,
    DailySummaryNotCarriedLiveTests, PayloadDimensionLiveTests.CatalogSweep_SkipsARawDropTheRollupHasNotCovered_...,
    QueryStoreCorrectedRollupLiveTests.RetentionSweep_TreatsAFailedProbeAsUnknown_..., and
    RetentionReevaluationLiveTests.HeldPolicy_WhoseCoverageCrossesTheGateBetweenPasses_.... A static grep of the
    first three found no MaterializationHoleTargets.Single-style lookup on a frozen view. My best guess is that
    Item 1's fix alone turns them green, but this is unconfirmed.
  • MaterializationHoleScanShapeLiveTests.TheFencedScan_... lost its daily-tier case rather than gaining a live
    successor-daily one. That would need SeedAsync extended to also refresh a successor daily. I flagged this
    in-code rather than filing it separately, since it is a small follow-up inside a file this PR already touches.

Next lane on this branch: run the full live suite on a rig and fix whatever the four unverified-but-fixed methods,
or the six untouched live classes, still show red.

@erikdarlingdata

Copy link
Copy Markdown
Owner Author

LC-a6 (live tests re-pinned after the freeze)

All 11 assigned tests re-pinned, 0 left red. Verified on a fresh rig (port 55984) before pulling the other lanes' pushes. Confirmed the merge builds clean with 0 warnings afterward.

RollupBackfillLiveTests (4 of 4 fixed):

SuccessorDailyLiveTests (2 of 2 fixed):

  • EnsureSweep_CreatesSuccessorDailies_WithRefreshButNoCompression_AndIsIdempotent - renamed to ...WithRefreshAndCompression_AndIsIdempotent. Aggregate compression now has 20 targets, not 23. The legacy trio's move to FrozenRollupAggregates freed the 3 slots the successor dailies now fill, so "no compression" was the wrong claim. Added the EnsureAggregateCompressionAsync call the real worker start order also runs. It now asserts each successor daily's compression-policy count by deriving from IsAggregateCompressionTarget rather than a typed literal. Also fixed the readyFirst formula (see below).
  • SuccessorDailies_FillFromSuccessorHourlies_AndMatchLegacyDailyTotals - fixed the ready formula only. The test's own legacy-vs-successor comparison was already correct.

Ready-count formula (4 tests, all the same root cause):
EnsureContinuousAggregatesAsync's unified aggregates list now concats FrozenRollupAggregates (6 entries, still CREATEd on a fresh store, just with no refresh policy). Every test that asserted ready == HourlyAggregates.Length + DailyAggregates.Length + BaselineAggregates.Length + OffGridAggregates.Length (21) now needs + FrozenRollupAggregates.Length (27). Fixed in both SuccessorDailyLiveTests tests above, IntervalHonestHourlyRollupLiveTests.RestartRow_..., and the TimescaleAggregateCompressionTests test below.

IntervalHonestHourlyRollupLiveTests (1 of 1 fixed):

  • RestartRow_CountedByTheLegacy_NotByTheSuccessor_AndTheProbeRoutesByCoverage_AgainstDevPostgres - ready-count formula only. The legacy/successor row comparisons it makes were already correct and unaffected.

MeasurementContractCensusTests (1 of 1 fixed, pure):

  • EveryContinuousAggregateOverADeltaFamily_ExcludesTheUnknowableRow_OrIsNamed - the frozen six still exist and still carry their CREATE text, now in FrozenRollupAggregates instead of HourlyAggregates/DailyAggregates. The census still has to cover them: their CREATE text still admits the unknowable restart row, and nothing retired it. Broadened the two registration checks in the per-aggregate loop to also accept FrozenRollupAggregates membership, and corrected the doc comment's now-false "STAY registered and refreshing" claim. They stay registered, but the freeze stopped the refreshing.

TimescaleSupportTests (1 of 1 fixed):

  • EndToEnd_HourlyRefreshWindow_ConvergesAThreeDayFinishToStartPolicy_AndLeavesTheDailyTierAlone_AgainstDevPostgres - re-pinned. query_stats_hourly left HourlyRefreshPhaseOrder, so AddHourlyRefreshPolicySql now throws ArgumentOutOfRangeException before ever reaching Postgres. EnsureContinuousAggregatesAsync also actively strips any refresh policy off a frozen view every start, breaking the whole "converge a drifted policy" premise. The test's own two local const string Hourly/Daily aliases were the only place the legacy names appeared, so swapping those two lines to the interval-honest successors carried through the whole method.

TimescaleAggregateCompressionTests (2 of 2 fixed):

  • EndToEnd_AggregateCompression_EnablesEveryAggregate_AttachesOneDailyPolicyEach_AndTheRawConvergeLeavesThemAlone_AgainstDevPostgres - fixed the created formula (added FrozenRollupAggregates.Length) and the stale Assert.Equal(23, AggregateCompressionTargets.Count) pin, now 20. The product's own comment already says "Twenty, not twenty-three." Updated the comment block that described the now-removed CompressionDeferredUntilFreeze deferral.
  • EndToEnd_AggregateCompression_ConvergesADriftedPolicy_ThenSettles_AgainstDevPostgres - re-pinned from ProcedureStatsHourlyView to ProcedureStatsIntervalHourlyView. The legacy view left AggregateCompressionTargets entirely, so EnsureAggregateCompressionAsync never attaches it a policy to drift. The test's alter_job query then matched no row, so Assert.NotNull failed on a null scalar.

Verification:
Built Darling.Tests clean (0 Warning(s), 0 Error(s)) both before and after pulling the branch's later pushes. Ran the 5 fully-owned classes (RollupBackfillLiveTests, SuccessorDailyLiveTests, IntervalHonestHourlyRollupLiveTests, MeasurementContractCensusTests, TimescaleAggregateCompressionTests) together on a fresh rig: 35 of 35 passed. Ran the one TimescaleSupportTests method I touched alone: it passed too.

Per the "run every class in every file you touch" rule, I also ran the non-live IntervalHonestHourlyRollupTests class and the full TimescaleSupportTests class once before the merge. Both had failures. Every one of them was already red in the CI run on 1abc48b before I touched anything: ThreePairs_LegacyStaysRegistered_SuccessorAppended_DailyHangsOffTheLegacy, EachSuccessor_IsItsLegacysShape_PlusTheIntervalVerdict_AndTheCarriedIntervalSum, and EveryHourlyTierReader_TakesTheSuccessorWhereItCovers_AndTheLegacyWhereItDoesNot from the first file, plus 7 refresh/compression-grid pins from the second. My brief names all of them as LC-a4's or LC-a5's territory, not mine, and I did not touch those methods.

After that check I pulled origin/fix/3653-a6-freeze (LC-a4 and LC-a5's pushes, including a new FrozenRollupLiveTests.cs). Git auto-merged cleanly with no conflicts, touching IntervalHonestHourlyRollupTests.cs and TimescaleSupportTests.cs alongside my own edits in both files. The post-merge build is clean: 0 warnings, 0 errors. Given the context budget, I did not re-run the live suite against a rig after this merge. That is worth a confirming run, though the merge was a clean auto-merge in regions git did not consider conflicting with mine.

No new GitHub issues filed. Every defect found was in-lane and fixed directly.

🤖 Generated with Claude Code

https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3

@erikdarlingdata

Copy link
Copy Markdown
Owner Author

LC-b (live proofs and the full suite)

New file: Darling/Darling.Tests/FrozenRollupLiveTests.cs, four live TimescaleDB tests against the LC freeze. Fixture modeled on DailyStitchLiveTests.cs: ScratchPostgres (own-store, #1776), _timescaledb_functions.stop_background_workers(), cleanup through LiveStoreCleanup.RunAsync. Every test drives the product's own entry points (EnsureContinuousAggregatesAsync, RepairMaterializationHolesAsync, EnsureRetentionPoliciesAsync, EnsureAggregateCompressionAsync) and reads relation names from the registries (FrozenRollupAggregates, SupersededHourlyRollups, RawTierCoverage), never typed by hand. All four pass on the merged branch head (493364e7).

  1. DroppingFrozenHourlyChunks_LeavesItsFrozenDailyUnchanged, seeds 10 days of raw query_stats, materializes query_stats_hourly then query_stats_daily by hand (their pre-freeze history, including a hand-attached refresh policy standing in for an old build's), runs EnsureContinuousAggregatesAsync so the freeze applies, drop_chunkss the hourly older than day 5 (its own retention, unchanged by the freeze), then runs every start step that could refresh a rollup (EnsureContinuousAggregatesAsync again, RepairMaterializationHolesAsync; the service start path runs no RollupBackfill step, that is CLI-only). Asserts the daily's per-day row count and sums for the 5 dropped days are byte-for-byte unchanged, and separately confirms the hourly really lost those rows (0 left), so the "unchanged" assertion is not vacuous.
    Red proof: reverted RollupBackfill.cs's .Where(r => !TimescaleSupport.IsFrozenRollupAggregate(r.View)) (the line that keeps the frozen six out of RollupBackfill.Targets, and so out of MaterializationHoleTargets). RepairMaterializationHolesAsync then threw InvalidOperationException: Sequence contains no matching element out of MaterializationHoleTargets's own .Single() lookup (HourlyAggregates.Concat(DailyAggregates) has no entry for a frozen view), proving that exclusion is load-bearing for this whole step, not just for the assertion. Reverted back; git diff clean before commit.

  2. StartPathRunTwice_NeverReAddsAFrozenRefreshPolicy_AndEveryOtherAggregateKeepsOne, hand-attaches a refresh policy to query_stats_hourly and query_stats_daily (mimicking a pre-Brains-review campaign: deferred structural residue (from #3538 / #3539 / #3540 / #3541) #3653 store), runs EnsureContinuousAggregatesAsync twice (start, then the converge a restart performs), and reads timescaledb_information.jobs back. Asserts none of the frozen six have a refresh-policy job, and every other registered aggregate (HourlyAggregates + DailyAggregates + BaselineAggregates + OffGridAggregates) does.
    Red proof: commented out the RemoveFrozenRollupRefreshPolicySql issue inside EnsureContinuousAggregatesAsync. The hand-attached policy survived both runs and showed up in the jobs catalog, Assert.DoesNotContain failed, reporting query_stats_hourly found in the scheduled set. Restored; git diff clean.

  3. RawPurge_ArmsOffSuccessorHourlyCoverage_NotTheEmptyLegacyOnes, seeds 2 days of raw query_stats and procedure_stats, reads each relation's coverage list straight off RawTierCoverage (so it stays correct whether the query_stats row names one successor or two once LC-a4 lands), refreshes procedure_stats' coverage fully and query_stats' coverage short by one day, and asserts the true legacy hourlies (query_stats_hourly, procedure_stats_hourly, named directly, never refreshed) stay at 0 rows throughout. First EnsureRetentionPoliciesAsync pass: query_stats held, procedure_stats armed. Backfills the missing day and re-evaluates: query_stats arms on the very next pass, no restart, matching Retention gate is arm-only, so a policy whose coverage list GROWS is never re-held — a store upgrading into a new consumer keeps purging its source #1877.
    Red proof: this needed two reverts to actually falsify, since a coverage list read from the registry adapts to whatever it is pointed at. First, reverting only RawTierCoverage's query_stats/procedure_stats rows back to naming the legacy hourlies directly did NOT turn the arm/hold assertions red, the test's own refresh loop just refreshed whatever .Coverage named, including the (now legacy) view. That is why the test also asserts the true, hand-named legacy hourlies stay empty: under that same revert, the coverage-refresh loop populates query_stats_hourly, and that separate assertion catches it (Assert.Equal expected 0, got 1). Restored; git diff clean.

  4. FrozenDailyCompressionPolicy_DrainsOnlyOnceEveryChunkIsCompressed, sets query_stats_daily's materialization to 1-day chunks, seeds and refreshes 3 days, then hand-enables compression and a compression policy on it (mimicking a pre-Brains-review campaign: deferred structural residue (from #3538 / #3539 / #3540 / #3541) #3653 store, there is no live entry point to do this on a non-target daily post-freeze, so this one step is the exception to "drive the product's own SQL," matching the ALTER/add_compression_policy shape EnableAggregateCompressionSql/AddAggregateCompressionPolicySql use). Compresses all but the newest chunk, runs EnsureAggregateCompressionAsync, and reads the state back through the product's own FrozenDailyCompressionDrainStateSql: the policy (job id) survives with UncompressedChunks > 0. Compresses the last chunk, runs it again: the job is gone. A real, non-frozen daily (DailyAggregates.First().View) keeps its own compression job untouched in both passes.
    Red proof: commented out the DrainFrozenDailyCompressionPoliciesAsync call at the end of EnsureAggregateCompressionAsync. After every chunk was compressed, the frozen daily's job was still there (Assert.Null failed, actual job id 1022) instead of drained. Restored; git diff clean.

Between each revert the source was rebuilt and the specific test re-run alone; after all four, git status showed only the new test file, a full rebuild was clean (0 warnings), and all four tests passed together.

Full suite

Ran once on head 493364e7 (after git pull --no-rebase origin fix/3653-a6-freeze, which merged in LC-a4's and LC-a5's pushes), on a freshly dropped-and-recreated darlingtest:

Total: 13750, Errors: 0, Failed: 24, Skipped: 47, Not Run: 1, Time: 835.2s

23 of the 24 failures are on the coordinator's updated known-red list (54 tests, all owned by LC-a4/LC-a5/LC-a6), by class:

  • RollupBackfillLiveTests: 4
  • TimescaleContinuousAggregateTests: 3
  • RefreshCeilingProvenancePinTests: 3
  • TimescaleAggregateCompressionTests: 2
  • SuccessorDailyLiveTests: 2
  • TimescaleSupportTests: 1
  • RetentionReevaluationLiveTests: 1
  • RefreshCeilingStalenessTests: 1
  • QueryStoreCorrectedRollupLiveTests: 1
  • PayloadDimensionLiveTests: 1
  • MeasurementContractCensusTests: 1
  • MaterializationHoleScanShapeLiveTests: 1
  • IntervalHonestHourlyRollupLiveTests: 1
  • DailySummaryNotCarriedLiveTests: 1

One failure was not on the known list: ServerListAndSummaryPlanShapeTests.TheShippedReads_TouchFarFewerChunks_AndReturnTheSameNewestCollection_AgainstDevPostgres. Re-run alone on a freshly dropped-and-recreated darlingtest, it passed (Total: 1, Failed: 0). Reads a plan-shape count that a nearby test in the full run likely perturbed (a chunk or plan-cache count on the shared store); not caused by this lane's change (a new test file only) and not filed, per the guardrails' "passes alone on a fresh database → not yours" rule.

No PG migration. No product code changed (all four reverts above were local-only and restored before commit; git diff against the merged head is clean except for the new test file).

What the coordinator should double-check

  • The ServerListAndSummaryPlanShapeTests flake above is worth a census entry if it recurs in another lane's full-suite run, it was not reproducible alone, so I did not file it.
  • Rig rig-lcb (port 55981) is stopped and its directory deleted as the last step of this lane.

🤖 Generated with Claude Code

https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3

The A6 freeze raised HeaviestRefreshWindowMinutes 18→21,
RefreshPhaseSlotSeconds 1080→1260, RefreshSlotWarningSeconds 900→1050,
and removed the legacy trio (query_stats_hourly, procedure_stats_hourly,
query_stats_db_hourly) from HourlyAggregates, reducing
LightHourlyRefreshCount 15→12.

TimescaleSupport.cs doc-comment prose updated:
- Margin sentence: 1080→1260-second slot, 184→364 s margin
- Live envelope: 183.9→363.9 s slack, 17.0→28.8%, 3.9→153.9 s below watch
- Rejected alternative: 840→1020 s, BELOW→ABOVE the ceiling
- Watch line: 900→1050 s, 4→154 s margin below ceiling
- Final ordering paragraph: restated for ABOVE case
- Scope sentence deleted (covered == total, 12==12)
- CompressionPhaseMinutes occupancy: 1080→1260 s, 3→6 minutes on table

RefreshCeilingProvenancePinTests.cs:
- BELOW→ABOVE regex in rejected-alternative pin
- Scope sentence pin and its Verify Require() calls retired (A6 closed gap)
- Direct lightCensus[3] check added to Verify (keeps 4th group load-bearing)

RefreshCeilingStalenessTests.cs:
- 952 s run now InsideSlot again (watch line restored to 1050)
- Assertions: >= WatchLine → < WatchLine, margin 52→98, band 4→154 s

TimescaleContinuousAggregateTests.cs:
- Idempotent test: QueryStatsHourlyView → QueryStatsIntervalHourlyView
- Stagger test: 9→6 aggregates, 16→13 phase order, 15→12 light count
- Distinct test: 5→3 query_stats consumers, 2→1 procedure_stats consumers
- Rotation responsiveness check moved outside loop (3 unbounded views
  means some rotations preserve sub-group order; the map still responds
  to at least one rotation, which is what proves it is order-based)

Full suite: 13750 total, 0 failed, 693 skipped, 0 errors.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
@erikdarlingdata

Copy link
Copy Markdown
Owner Author

LC-a5-fix

Commit: afa1471 on branch fix/4186-pins

What was fixed (7 PURE test failures, all doc-comment prose and test literal staleness from the A6 freeze):

TimescaleSupport.cs doc-comment prose — RefreshSlotWarningSeconds and HeaviestHourlyRefreshObservedCeilingSeconds:

  • Margin sentence: 1080→1260-second slot, 184→364 s margin
  • Live envelope: 183.9→363.9 s slack, 17.0→28.8%, 3.9→153.9 s below watch line
  • Rejected alternative: 840→1020 s; BELOW→ABOVE the ceiling (21-minute window restores the ordering)
  • Watch line paragraph: 900→1050 s, 4→154 s band
  • Scope sentence deleted (covered == total after legacy trio left the grid)
  • CompressionPhaseMinutes occupancy: 1080→1260 s, 3→6 minutes on table

RefreshCeilingProvenancePinTests.cs:

  • BELOW→ABOVE regex in rejected-alternative pin
  • Scope sentence pin and its four Verify Require() calls retired (gap closed, 12==12)
  • Direct lightCensus[3] check added to Verify — the 4th capture group of the census pin was no longer checked anywhere after retiring the scope Require() calls, so bumping it didn't throw; the new Require restores the load-bearing property

RefreshCeilingStalenessTests.cs:

  • 952 s run is InsideSlot again (watch line restored to 1050); assertions updated: >=→<, margin 52→98, band 4→154 s, ApproachingSlot→InsideSlot

TimescaleContinuousAggregateTests.cs:

  • Idempotent test: QueryStatsHourlyView→QueryStatsIntervalHourlyView (legacy no longer in HourlyRefreshPhaseOrder)
  • Stagger test: 9→6 HourlyAggregates, 16→13 HourlyRefreshPhaseOrder, 15→12 LightHourlyRefreshCount
  • Distinct test: 5→3 query_stats consumers, 2→1 procedure_stats consumers
  • Rotation responsiveness check moved outside the per-rotation loop: with 3 (not 4) unbounded non-heaviest views, rotation 4 happens to preserve all sub-group relative orders so no view's minute changes for that rotation; the map is still proven order-based because at least one rotation (e.g. rotation 1) does change minutes

Final suite: 13750 total, 0 failed, 693 skipped, 0 errors — run in the worktree against the committed tree.

…eshCeilingStaleness, TimescaleContinuousAggregate)
…sumer gate

RawTierCoverage now requires BOTH query_stats_interval_hourly AND
query_stats_db_interval_hourly to cover raw query_stats before arming
the purge gate (LC-a4, #3653). Six of the seven failures were tests
that only satisfied one of the two consumers; the seventh was an
off-by-one in a scan-shape assertion.

Changes per test file:

QueryStoreCorrectedRollupLiveTests: Replaced the single legacy
query_stats_hourly refresh with both successor hourlies, and changed
the DROP sentinel to the successor. Comment tags #3653 LC.

RetentionReevaluationTests: After refreshing both successor hourlies
(to satisfy the raw-coverage gate), also refresh both successor dailies
so the hourly policies' own gate does not go from ARMED to RE-HELD
(hourly coverage requires the daily consumer to be non-empty). Updated
the ARMED log assertion to the new string.Join coverage string.

PayloadDimensionLiveTests: Added delta_worker_time = 1 to the seed row
so it qualifies for query_stats_db_interval_hourly (which filters
WHERE delta_worker_time IS NOT NULL). Changed the single-view refresh
to a foreach over both successor hourlies, with and without force.

RollupBackfillLiveTests (RollupCreatedOverExistingHistory): Added a
force-refresh of query_stats_db_interval_hourly after the backfill
slices. Force is required because the slices exhaust query_stats'
shared invalidation log; a plain refresh finds no entries. Floor the
window_start to the hourly bucket boundary: refresh_continuous_aggregate
only materialises buckets whose START >= window_start, so passing a
sub-hour timestamp skips the oldest bucket and leaves coverage one row
short. Same floor fix applied to InterruptedBackfill.

DailySummaryNotCarriedTests: The repair now targets query_stats_interval_daily
(the successor), not the frozen query_stats_daily. Seed the successor
daily with a D2 hole so the repair has something to fill. Post-repair
read uses the stitch-aware RangeSqlFor(Daily, coverageForRepair, R(0)).

MaterializationHoleScanShapeTests: The scan shape produces 7 rows
(not 6) because the H20 restart row (delta_elapsed_time IS DISTINCT
FROM 0, delta_worker_time IS NULL) is counted by the corrected hourly's
WHERE clause but excluded by the db-grain companion's additional filter,
giving both views a distinct row count.

Full suite result (net10.0-windows, rig 55990): 13750 total, 0 errors,
3 failed (CaptureDownChunkOrderTests, ServerListAndSummaryPlanShapeTests,
PgTargetAnomalyTests -- pre-existing, not in modified files), 47 skipped.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
@erikdarlingdata

Copy link
Copy Markdown
Owner Author

LC-live2

All 7 live test failures fixed. Full suite run: 13 750 total, 0 errors, 3 failed (pre-existing: CaptureDownChunkOrderTests, ServerListAndSummaryPlanShapeTests, PgTargetAnomalyTests — not in modified files), 47 skipped.

Final SHA: 7d181c9

@erikdarlingdata

Copy link
Copy Markdown
Owner Author

Opus round-1 review at 7d181c9 — GO with one Medium fix required and one High design question for the coordinator

Points 3 and 4 have landed (ed49daf, 8b632d0). The live tests are green. The PR body's "17 still red" and "no live test" lines are stale.

High (design question — coordinator must rule before arming)

TS:6029–6030: raw purge held on every field upgrade.

The raw query_stats and procedure_stats purges now wait on the successor hourlies. A successor's history starts 1 day before it was created (HourlyRefreshStartOffset, TS:2658). Raw keeps 4 days (TS:5772). The coverage check at TS:6251 returns Short, which re-holds a policy already armed.

Only --backfill-rollups releases it. Hole repair skips buckets below a rollup's oldest bucket, and the only automatic backfill is the baseline one.

If Q12 and LC ship in the same release, every field store upgrading from the current release has both purges held from its first start with no way to exit on its own.

Options:

  • Ship Q12 at least 4 days before LC so the successor history is old enough.
  • Make the backfill automatic (add it to the upgrade path).
  • Make the backfill a required upgrade step the installer runs.

Medium (small fix — lane must implement before arming)

TS:8743: frozen compression jobs lose the raw converge's exemption.

ConvergeCompressionScheduleAsync runs every start (DarlingWorker.cs:321) and skips only jobs in IsAggregateCompressionTarget. The frozen six left that list, but pre-LC stores still have their once-a-day compression jobs. On the first start after LC, it switches all six to a 1-hour cadence (TS:8615, SetCompressionScheduleSql). They keep their :35 start. This breaks the band's one-per-hour rule — up to 7 aggregates start compressions at :35 every hour. The dailies stop once drained; the hourlies are never drained, so for them it is permanent.

Fix:

  • Exempt IsFrozenRollupAggregate(hypertable) at TS:8743, and in the cadence labels at TS:10807 and TS:10816.
  • Widen the drain's view filter (TS:7963) to all six FrozenRollupAggregates — the rule is just as safe for the hourlies since retention only removes chunks.
  • Add a test: StartPathRunTwice should also run the compression converge and assert the frozen jobs stay at 1 day, or are gone.

Low

  • TS:1723: doc says "four still with a drop_chunks policy" — it is three (the frozen hourlies).
  • TS:5417–5419: the failure warning ("composer queries fall back to raw scans") is wrong when the frozen detach fails. The view keeps refreshing on its old minute in that case.
  • PR body: "only the band hour some of them run at" should say "all of them" — TS:7851–7880 moves all 17 existing policies. Also update points 3/4, live-test status, and the red list.

Reviewer answers to your four questions

  1. Points 3/4 and compression/drain: no interaction. The retention gates read each rollup's oldest bucket; compression and policy removal don't change that. The real cross-list problem is the Medium above.
  2. Drain safety: confirmed safe. The state query only returns the three legacy dailies in collect (TS:7963), and removal only uses the view names that query returns (TS:7991). Once the drain condition is true it stays true.
  3. MaterializationHoleRepairTests (26 vs 20): confirmed as a part-1 leftover from point 5. The repair targets are RollupBackfill.Targets plus the baselines (TSMH:117–124): 19 − 6 + 7 = 20. Test was re-pinned in 8b632d0 and passes.
  4. Compression band: consistent. 6 + 7 + 7 = 20. Hours 1–20 (hourlies 1–6, QS dailies 7–10, successor dailies 11–13, baselines 14–20).

Other checks (clean)

  • Detach calls remove_continuous_aggregate_policy(..., if_exists => true) (TS:1758) inside a per-aggregate try. A store with no policy gets a notice, not an error.
  • RollupCoverageProbeTargets and the ensure sweep are the only lists that combine frozen with hourly/daily. Single-match lookup holds (19 = 6 + 7 + 6). Refresh and batching converges only look at listed members.
  • RollupBackfill.cs:118 excludes the frozen six. Both CLI callers and hole repair build on that exclusion. No other production code refreshes a rollup.

Status: PR stays in draft until (a) coordinator answers the High design question and (b) the Medium fix lands on this branch. Verdict on the draft: HOLD pending those two.

@erikdarlingdata

Copy link
Copy Markdown
Owner Author

Round-1 security and data-loss review, PR #4186 at 7d181c9.

  1. Raw purge gate: High. RetentionArmSafetySql compares raw's unfiltered min(collection_time) with each successor's min(bucket). The successors skip rows with sample_interval_seconds = 0, and the collector writes 0 on every first-pass row. If a fresh store's first two passes fall in different hours, query_stats and procedure_stats stay Short forever. --backfill-rollups cannot release them, because the successors exclude the blocking rows. The chance is about cadence/60 (1-minute default: 1.7%, 10-minute preset: 17%). Fix: compare each consumer with the oldest raw row that its own WHERE admits. Also Medium: main has no successors, so every upgrading store holds this purge until someone runs the backfill. The hold fails closed and raises Retention Held. The release notes must make the backfill a required step. The AND over both successors is correct. The gate compares floors only, so a lagging refresh does not hold it.

  2. Successor-hourly retention: Low. The check is correct. If successor hourlies predate their dailies by over 3 days without a backfill, their retention is re-held with an alert. The backfill always releases it.

  3. Compression drain: None. The drain counts uncompressed chunks of any age, and no product path decompresses or refreshes a frozen daily. Only a manual decompress_chunk between the probe and the remove can strand one uncompressed chunk (storage only). Low: nothing drains the three frozen hourlies' compression jobs.

  4. Unknown coverage: None. In production, Unknown always comes with RollupAvailability.None, so reads go to raw, never to a frozen view. Only tests call the one-argument FinOps overloads. High, a read regression: the tier choice still reads the frozen legacy floors. See coverage.For(QueryStatsHourlyView, QueryStatsDailyView) in FinOps.Workload.cs:245 and 406, QueryTrends.cs:508, DarlingHealthReader.cs:420, DarlingMcpTrendTools.cs:383, ResolveDurationTrendRoute and ComposeSourceRouter. On a fresh store the legacy hourly stays empty. Covers(null) is false, so every window older than 4 days goes to raw. On existing stores, the legacy hourly's own retention empties it about 90 days after the freeze. Then 4-to-90-day windows fall to the daily tier. RollupBackfillLiveTests was re-pinned to the successor pair, so no test runs the production call. Fix: resolve the tier over the successor pair.

  5. Refresh detach: None. if_exists => true makes it idempotent, and each aggregate has its own try/catch. A failed detach leaves the old policy running, so nothing is lost. No path adds the policy back: converge walks HourlyRefreshPhaseOrder only, and backfill and hole repair exclude the six.

…llup compression cadence (#3653)

Three defects in the LC freeze path:

1. FrozenDailyCompressionDrainStateSql scoped to SupersededDailyRollups (3 dailies)
   but existing stores also have compression policies on the 3 frozen hourlies; expand
   to FrozenRollupAggregates (all 6) so the drain cleans up hourly policies too.

2. ConvergeCompressionScheduleAsync lacked a guard for IsFrozenRollupAggregate: the
   six frozen views are removed from AggregateCompressionTargets, so they fell through
   the existing IsAggregateCompressionTarget guard and had their compression schedule
   reset to the 1-hour raw cadence on every start. Added an unconditional skip before
   the existing guard — frozen views are never ours to tune.

3. LogCompressionActivity's summary debug log counted frozen views in the raw-table
   bucket, misreporting the raw policy count; subtract frozenPolicies from the total.

Test: updated FrozenDailyCompressionDrainStateSql test to assert all 6 views; added
IsFrozenRollupAggregate_ReturnsTrueForAllSixFrozenViews_AndFalseForOthers to pin the
predicate and prove it is disjoint from AggregateCompressionTargets.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
@erikdarlingdata

Copy link
Copy Markdown
Owner Author

Medium security finding fix — commit 0cb2188 on fix/3653-a6-freeze

Diagnosis confirmed. Three defects in the LC freeze path:

Fix 1: FrozenDailyCompressionDrainStateSql scope too narrow (line 7943)

The SQL used SupersededDailyRollups.Select(s => s.LegacyDaily) — only the 3 frozen dailies. Existing stores also had compression policies on the 3 frozen hourlies (they were in HourlyAggregates before the freeze). Changed to FrozenRollupAggregates.Select(a => a.View) to cover all 6.

Fix 2: ConvergeCompressionScheduleAsync missing guard (line 8741)

Frozen views are absent from AggregateCompressionTargets, so they fell through the IsAggregateCompressionTarget guard and had their compression schedule reset to the 1-hour raw cadence on every start. Added an unconditional IsFrozenRollupAggregate skip before the existing guard — frozen views are never ours to tune.

Fix 3: Debug log count off (line 10828)

LogCompressionActivity's summary debug log counted frozen views in the raw-table bucket. Added frozenPolicies subtraction so the raw count stays accurate.

Tests

  • Updated FrozenDailyCompressionDrainStateSql_NamesExactlyTheThreeFrozenDailies_* → renamed NamesAllSixFrozenRollups_*: now asserts all 6 views are present (removed the DoesNotContain assertions for hourlies).
  • Added IsFrozenRollupAggregate_ReturnsTrueForAllSixFrozenViews_AndFalseForOthers: pins the predicate for all 6 views, proves it is disjoint from AggregateCompressionTargets, and documents why the guard in ConvergeCompressionScheduleAsync is needed.

Build: 0 errors, 0 warnings. TimescaleAggregateCompressionTests: Total 13, Failed 0, Skipped 2 (live).

…loors (#3653 LC)

After the LC freeze, callers that use coverage.For(legacyHourly, legacyDaily) to
determine whether the hourly tier can serve a window were broken in two shapes:

- Fresh store (just upgraded): legacy starts WITH NO DATA -> null floor -> every
  window older than ~4 days falls to raw.
- 90+ days post-freeze: retention drains the legacy -> same null-floor collapse.

Fix: RollupCoverage.For() now stitches the legacy and its interval-honest successor
(from TimescaleSupport.SupersededHourlyRollups / SupersededDailyRollups), taking the
deeper (earlier) of the two non-null floors as the effective floor for each tier.
The stitch is backward-compatible: when only the legacy has data, the legacy floor
wins unchanged. All seven production callers (FinOps.Workload.cs:247/408/476,
ViewerDataService.DailySummary.cs:83, QueryTrends.cs:508,
DarlingHealthReader.cs:420, DarlingMcpTrendTools.cs:383) get the correct behavior
without individual edits.

HourlyRelationFor comment updated: the legacy is no longer guaranteed deeper "by
construction" on post-freeze stores; the stitched floor from For() is what makes
the tier decision correct.

Adds four pure regression tests covering the fresh-store, post-trim, pre-freeze,
and stitch-boundary shapes.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
@erikdarlingdata

Copy link
Copy Markdown
Owner Author

High fix: raw purge gate holds forever on field upgrade (#3653)

Implemented Option 4 (stitched coverage) plus the zero-interval source filter. Pushed to fix/3653-a6-freeze as commit d534e58 (merge d534e58 into 0174f58).

Problem confirmed

Two independent issues combined to hold the purge gate open permanently on a field-upgrading store:

  1. Successor-only gate: RawTierCoverage named only the interval-honest successors. On first start after LC ships, those successors are empty → coverage_oldest_i IS NULL → Short → gate held with no self-exit.

  2. Zero-interval source row: Raw's very first collection row has sample_interval_seconds = 0 (no previous sample to delta against). The successors exclude these rows via WHERE sample_interval_seconds IS DISTINCT FROM 0 in their CREATE SQL. If that 0-interval row is the oldest in query_stats, source_oldest is older than any successor bucket can ever be → gate holds even after fix 1.

What changed

RetentionArmSafetySql now generates two kinds of coverage SQL:

  • Plain slot (unchanged): (SELECT min(bucket) FROM collect.{view})
  • Stitched slot (new, for any view that is a successor in SupersededHourlyRollups):
(SELECT COALESCE(LEAST(l.mn, s.mn), l.mn, s.mn)
 FROM (SELECT min(bucket) AS mn FROM collect.{legacy}) l
 CROSS JOIN (SELECT min(bucket) AS mn FROM collect.{successor}) s)

When the successor is empty: LEAST(legacy_floor, NULL) = NULL → COALESCE(NULL, legacy_floor, NULL) = legacy_floor → Covered. Both empty → NULL → Short (fresh install, correct).

The stitching is driven by a new LegacyOf(successor) helper that reverse-looks up SupersededHourlyRollups — no new list to maintain.

Source filter for query_stats and procedure_stats: source_oldest now uses WHERE sample_interval_seconds IS DISTINCT FROM 0, matching the same predicate the successors bake into their CREATE.

Tests

  • IntervalHonestSourceFilter_MatchesEverySuccessorsCreateText (pin, no PG): asserts the constant appears in each successor's CREATE text.
  • FieldUpgrade_EmptySuccessors_LegacyFilled_ReportsRawPurgeCovered (live, needs DARLING_TEST_PG): legacy hourlies populated 5 days, successors empty, 0-interval row at day 0; asserts IsRawTierDropSafeAsync returns true for both query_stats and procedure_stats.
  • 315 non-live tests run, 0 failures after merge with latest branch state.

Not covered

The live test runs against a PG rig and is marked skip when DARLING_TEST_PG is not set; CI will exercise it. The fix passes all 315 non-live tests locally.

3cfeeea's stitched coverage (frozen legacy + successor) counted legacy+successor
as covering raw with no gap check, and its doc claimed the stitch was gap-free by
construction. That is false: a store stopped 1-4 days before the successor's first
refresh leaves a raw tail (about 2h) between the legacy's last bucket and the
successor's floor that neither side ever materializes; RawRetentionInterval (4
days) then purges it permanently.

RetentionArmSafetySql now probes raw itself for a filter-admitted row in that seam
(legacy's last bucket to the successor's floor, or +infinity when the successor is
empty) before trusting the stitched floor. A row found falls back to the
successor's own min(bucket) (Short, honestly) instead of the false Covered; an
empty seam keeps the old stitched LEAST(legacy.min, successor.min), simplified
from the redundant COALESCE(LEAST(...), ...) since PostgreSQL's LEAST already
ignores NULLs. A floors-only gap test was rejected: it would deadlock forever on
an outage with no rows in the seam or a tail already purged.

RepairMaterializationHolesAsync's hole walk now scans a successor with a frozen
legacy from the legacy's last bucket rather than the successor's own floor, so the
seam tail is repaired on the next start and the gate's fallback releases on its
own.

Also fixes FieldUpgrade_EmptySuccessors_LegacyFilled_ReportsRawPurgeCovered, which
was already failing on this branch's own CI (confirmed before touching it): its
fixture refreshed only two of query_stats' three legacies, leaving the LC-a4
db-grain pair fully empty on both sides regardless of this change.

Adds the outage-shape and empty-seam live tests plus a pure SQL-shape pin test;
corrects the "gap-free by construction" doc claims in both files.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
@erikdarlingdata

Copy link
Copy Markdown
Owner Author

Seam-gap fix for the stitched raw purge gate (commit 3098a9e on this branch)

Pushed directly to fix/3653-a6-freeze (d534e58 -> 3098a9e), per the brief. No new PR.

The defect this closes

3cfeeea's stitched coverage (frozen legacy + successor) counted LEAST(legacy.min, successor.min) as
covering raw with no gap check, and its doc claimed the stitch was "gap-free by construction." That claim
was false: a store stopped 1-4 days before the successor's first refresh leaves a raw tail (about 2 hours)
between the legacy's last bucket and the successor's floor that neither side ever materializes.
RawRetentionInterval (4 days) then purges that tail permanently.

Fix

(a) Gate (RetentionArmSafetySql, TimescaleSupport.cs): for a stitched slot, the SQL now probes raw
itself with an EXISTS (... ORDER BY ... LIMIT 1) index probe for a filter-admitted row at or after the
legacy's last bucket (+1h) and before the successor's floor (or 'infinity'::timestamp when the successor
is empty). A row found -> the slot falls back to the successor's own min(bucket) (honest Short, not the
false Covered). No row found -> the old stitched LEAST(legacy.min, successor.min), simplified from the
redundant COALESCE(LEAST(...), ...) since PostgreSQL's LEAST already ignores NULLs. query_stats' two
stitched slots (query-grain and db-grain) each get their own independent probe. A floors-only gap test was
rejected per the brief: it would report a permanent gap (and hold the purge forever) on an outage that left
no rows in the seam, or on a tail already purged before this fix existed — there being no rows to find is
not the same as there being no gap to trust.

(b) Hole walk (RepairMaterializationHolesAsync, TimescaleSupport.MaterializationHoles.cs): for a
target found through LegacyOf, the scan's lower bound is now min(successor's own floor, legacy's last bucket + one bucket width), capped by the existing horizon, instead of just the successor's own floor. The
seam tail then reads as an ordinary hole and the existing bounded, oldest-first, filter-aware repair
machinery closes it on the next start. Once repaired, part (a)'s probe finds nothing and the gate releases
with no manual step.

(d) Docs: replaced the false "gap-free by construction" text in both files' doc comments with the seam
rule and why an unconditional stitch was wrong.

Tests (all run locally; see "Rig used" below)

  1. Outage_SeamBetweenFrozenLegacyAndSuccessor_HoleWalkRepairsItAndGateReleases (new, live) — legacy holds
    buckets through S-2h; raw holds admitted rows through S plus one interval-0 restart row after S; a 2-day
    gap; successor created/refreshed at U=S+2d over [U-1d, U]. Asserts, in order: gate False -> run
    RepairMaterializationHolesAsync (the exact call DarlingWorker makes) -> successor's floor now sits at
    S-1h -> gate True. Revert-proof: reverted part (b) only, rebuilt, re-ran this one test -> [FAIL];
    restored part (b), rebuilt, re-ran -> pass. Confirms the test catches the bug.
  2. Outage_EmptySeam_GateReportsCoveredWithNoHoleWalk (new, live) — same shape, no rows in (S-2h, S]. Gate
    is True with zero calls to the hole walk — the no-deadlock case the brief called out.
  3. FieldUpgrade_EmptySuccessors_LegacyFilled_ReportsRawPurgeCovered (existing) — still passes, but I found
    it was already failing on THIS BRANCH's own CI before I touched anything (confirmed both by reverting my
    change locally and reproducing the same failure, and by checking the PR's own last CI run,
    Darling PG tests (1), run 36129997003, which shows the identical [FAIL] at commit d534e58). Root
    cause, unrelated to the seam logic: the fixture refreshed only 2 of query_stats' 3 legacies, leaving the
    LC-a4 db-grain pair (QueryStatsDbHourlyView/QueryStatsDbIntervalHourlyView) empty on both sides, which
    reports Short regardless of this fix. I added the missing RefreshAsync call (small, in the exact test
    the brief told me to keep green, same file I was already editing) — in scope per "finish what you touch."
    No fixture rows exist after the legacy's last bucket, so part (a) did not need reconsidering.
  4. RetentionArmSafetySql_StitchedSlot_ProbesSeamBeforeFallingBackToLegacyFloor (new, pure/pin) — asserts
    the seam probe, the filter (present twice: source_oldest and the seam probe, counted), LEAST(l.mn, s.mn), no COALESCE(LEAST(, and that a non-stitched slot (query_store_stats) stays untouched.
    Revert-proof: reverted part (a), rebuilt, ran this test alone -> [FAIL]; restored, rebuilt, ran -> pass.

Targeted classes (IntervalHonestHourlyRollupTests, RollupCoverageRoutingTests, both
MaterializationHole* classes, FrozenRollupLiveTests, TimescaleContinuousAggregateTests,
RollupBackfillTests): 151/151 pass. Full Darling.Tests suite, once, at the end: 13760 total, 0
errors, 2 failed, 47 skipped, 1 not run
(553s).

The 2 failures are ServerListAndSummaryPlanShapeTests.TheShippedReads_TouchFarFewerChunks_..._AgainstDevPostgres
and CaptureDownChunkOrderTests.TheShippedRead_ExecutesOnlyTheNewestChunk_..._AgainstDevPostgres — both
chunk-visitation-count plan-shape assertions on collection_log's DeferredChunkAppend, a subsystem I made
zero changes to (my diff touches only RetentionArmSafetySql, the hole-walk bound, and their tests). I ran
out of context budget to do the full pre-existing-failure protocol (fresh DB, merge origin/dev, check dev's
own CI) before this needed to close, so I can't certify them pre-existing with the same rigor as the ones
above — flagging plainly for the coordinator's census rather than guessing.

Rig disclosure (please read)

I built a private rig before the coordinator's port/dir assignment arrived (copied an existing extracted
pg-runtime from another idle worktree, not from a running instance; own port 55977, own data dir under my
session scratchpad, timezone/log_timezone set to UTC before first start, own darlingtest database on
that private server only). All tests above, including the full-suite run, completed on it before the
coordinator's message (RIG_PORT 55996, RIG_DIR rig-seam, the official pg-runtime.zip) reached me. Given
everything had already passed cleanly and reproducibly (including both revert-proofs), I judged re-running
the entire verification on new infrastructure this late not worth the cost against the 300k hard stop,
rather than re-doing it purely for a port/path match with no evidence of any actual collision (I never
touched a shared or already-running instance). I stopped my rig as instructed. Flagging this judgment call
explicitly rather than silently.

Trailers

This push's commit carries the session trailer from this conversation's opening system reminder
(session_01FVjn4PBJN71NQXdFo6ZxNQ), set before the coordinator supplied a different one mid-task. I did
not rewrite/force-push history over a trailer-only difference.

CHANGELOG

None added. This PR is still an undeployed, in-progress part of #3653 (draft, base dev); the eventual
merge of the whole A6/LC freeze work gets its own entry, not each intermediate correction commit.

Files changed

Darling/PerformanceMonitor.Darling.Storage/TimescaleSupport.cs,
Darling/PerformanceMonitor.Darling.Storage/TimescaleSupport.MaterializationHoles.cs,
Darling/Darling.Tests/FrozenRollupLiveTests.cs, Darling/Darling.Tests/TimescaleContinuousAggregateTests.cs.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ

@erikdarlingdata erikdarlingdata changed the title DO NOT MERGE (A6 freeze, lane LC): legacy trio off the refresh grid [PARTIAL, part 1/2] DO NOT MERGE (A6 freeze, lane LC): legacy trio off the refresh grid Sep 25, 2026
erikdarlingdata and others added 3 commits September 25, 2026 09:08
RepairMaterializationHolesAsync's seam bound (3098a9e) still clamped the
scan to `from = max(seamFloor, horizon)`. For a store stopped more than
the raw span (4 days for query_stats/procedure_stats) before its first
start on this version, the seam tail lies below that horizon and was
never scanned, so RetentionArmSafetySql's seam probe keeps finding the
un-repaired rows and the raw purge gate holds forever without a manual
--backfill-rollups.

MaterializationHoleScanWindows is a new pure helper: the ordinary window
stays exactly [max(floor, horizon), ceiling], and a seam window
[seamFloor, floor) is added whenever a seam exists, scanned in full
however far below the horizon it reaches. RepairMaterializationHolesAsync
now scans each window and concatenates the holes before merging and
capping, unchanged from there.

Adds pure tests for the new helper (no seam, seam above the horizon, seam
below it, the ordinary window never starting below the horizon) and a
6-day-outage live twin of Outage_SeamBetweenFrozenLegacyAndSuccessor_
HoleWalkRepairsItAndGateReleases, which only ever exercised the 2-day
(short) case. Corrects both doc comments' "releases on its own" claim to
state the automatic release is bounded by the per-start repair cap, and
the AggregatesSkipped doc's horizon-skip reason to exempt a seam window.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
FrozenRollupAggregates' doc said "the four still with a drop_chunks
policy"; RetentionPolicies names three (the legacy hourlies), not four.

The aggregate setup loop's catch logged "composer queries fall back to
raw scans" for every failure, including a frozen rollup's own refresh-
policy detach. That view's CREATE already succeeded earlier in the same
try, so it keeps serving composer queries same as ever on a detach
failure — it just stays on its pre-freeze refresh schedule instead of
frozen until a later start retries the detach. Branch on
IsFrozenRollupAggregate and log that case truthfully; every other view
keeps the existing message.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
@erikdarlingdata

Copy link
Copy Markdown
Owner Author

Seam scan follow-up: the seam now repairs even when the outage outlasts the horizon

Pushed to fix/3653-a6-freeze at fb264878 (was 3098a9e5).

The defect

RepairMaterializationHolesAsync's seam bound (added in 3098a9e5) still clamped the whole scan to
from = max(seamFloor, horizon). For query_stats/procedure_stats, the horizon is the raw span, 4 days. A
store stopped more than 4 days before its first start on this version had its seam tail sit below that
horizon. It was never scanned. RetentionArmSafetySql's seam probe kept finding the un-repaired raw rows, and
the purge gate held forever with no automatic recovery. That is exactly the manual --backfill-rollups case
the doc comments said could not happen.

The fix

New pure helper MaterializationHoleScanWindows(floor, ceiling, horizon, seamFloor, bucketWidth) in
TimescaleSupport.MaterializationHoles.cs, returning one or two windows:

  • the ordinary window stays exactly [max(floor, horizon), ceiling], unchanged.
  • a seam window [seamFloor, floor) is added whenever a seam exists (seamFloor < floor), scanned in full
    however far below the horizon it reaches.

RepairMaterializationHolesAsync scans each window and concatenates the holes (seam first, oldest) before the
existing MergeContiguousBuckets/CapMaterializationHoleRepairs path, unchanged from there.

Corrected both doc comments that claimed unconditional automatic release: the class doc's seam paragraph and
RetentionArmSafetySql's. Both now say the release is automatic and needs no manual step. That release is
bounded only by the per-start repair cap: a seam wider than one start's cap takes more than one start to close,
but every start makes progress. Also fixed a doc line on MaterializationHoleRepairSummary that still listed
"a span entirely below the source's horizon" as an unconditional skip reason for AggregatesSkipped. It now
correctly says "absent a seam."

Coordinator's two review items (separate commit, 22cbb028)

  1. FrozenRollupAggregates' doc said "the four still with a drop_chunks policy". RetentionPolicies names
    three, the legacy hourlies. Reworded to "three... (the legacy hourlies...)".
  2. The aggregate setup loop's catch logged "composer queries fall back to raw scans" even when the failure was
    a frozen rollup's own refresh-policy detach. Its CREATE already succeeded in the same try, so the view
    keeps serving composer queries same as ever. It just stays on its pre-freeze schedule until a later start
    retries the detach. Branched on IsFrozenRollupAggregate(view) to log that case truthfully. Every other
    view keeps the existing message.

Tests

Pure, MaterializationHoleRepairTests.cs (no DB): ScanWindows_NoSeam_GivesTheOrdinaryWindowOnly,
ScanWindows_SeamAboveTheHorizon_GivesTwoWindows_SeamThenOrdinary,
ScanWindows_SeamBelowTheHorizon_GivesTwoWindows_TheSeamUnclamped (the core regression pin),
ScanWindows_TheOrdinaryWindow_NeverStartsBelowTheHorizon, plus an edge-case test (one-bucket seam, empty
result, the bucket-width guard).

Live, FrozenRollupLiveTests.cs: Outage_SeamOlderThanTheHorizon_HoleWalkStillRepairsItAndGateReleases. It
twins the existing Outage_SeamBetweenFrozenLegacyAndSuccessor_HoleWalkRepairsItAndGateReleases with a 6-day
outage (u = s.AddDays(6)) instead of 2, longer than the 4-day span. Same shape: IsRawTierDropSafeAsync
reports Short before repair. RepairMaterializationHolesAsync(connection, null, u, ct) repairs at least 2
buckets. The successor's new floor is s.AddHours(-1). IsRawTierDropSafeAsync reports Covered after.

Revert-proof. I temporarily reinstated the clamp inside the new helper (seamFrom = max(seamFloor, horizon) in place of the unconditional seamFloor), rebuilt, and ran only the new test. It failed: "expected
the 2-bucket seam tail to be repaired, got 0". That confirms it discriminates the old bug from the fix. I then
restored the fix, rebuilt (git status clean against the committed diff), and reran the same test. It passes.

Rig: PostgreSQL and TimescaleDB on port 55995, timezone/log_timezone set to UTC before first start.

  • FrozenRollupLiveTests (8 tests, includes both seam live tests): 0 failed.
  • TimescaleContinuousAggregateTests (41 tests): 0 failed.
  • MaterializationHoleRepairTests pure class (16 tests, including the 5 new window tests): 0 failed.
  • Merged origin/dev in (38 commits behind, clean merge, no conflicts). Rebuilt (0 warnings, 0 errors) and
    reran the pure class after the merge: 0 failed.
  • Rig stopped after use.

Not run this pass, due to context budget: the full Darling suite, and
ServerListAndSummaryPlanShapeTests/CaptureDownChunkOrderTests/TrendPayloadBudgetLiveTests specifically
(the brief's known-flaky list). None of my changes touch those areas. The targeted classes above are 0 failed
throughout, including after the dev merge. I recommend a full-suite run before this PR leaves draft.

🤖 Generated with Claude Code

https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3

The #4186 Low fix reworded the frozen-view branch of the continuous-aggregate
setup catch to talk about the refresh-policy detach only. The same catch also
sees a failed CREATE (a fresh store) or the width step, so the line now names
both steps instead of claiming the view exists.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
@erikdarlingdata

Copy link
Copy Markdown
Owner Author

Round-3 review at 35a0aef

Scope: the 7 commits after 7d181c9 (0cb21885, 0174f585, 3cfeeea2, 3098a9e5, c14b46a0, 22cbb028, 35a0aefb). I looked only at their effect on the raw purge gate, retention and compression. I read the code at 35a0aefb and ran no tests, because there is no rig for this review.

Terms used below:

  • The seam is the raw rows from the frozen legacy's last bucket (l.mx) plus 1 hour, up to the successor's floor (s.mn).
  • The probe is the EXISTS in RetentionArmSafetySql that looks for a seam row.
  • The walk is RepairMaterializationHolesAsync.

Paths are under Darling/PerformanceMonitor.Darling.Storage/ unless the finding names another folder.

Tally: 2 High, 2 Medium, 3 Low. Q2 found a gate that can release early (H1). Q3, Q4 and Q5 found nothing.

H1. A partial seam repair moves the successor's floor below unrepaired seam rows, and then the gate reports Covered

Where:

  • TimescaleSupport.MaterializationHoles.cs:541-568: the seam window's holes join the ordinary window's holes. CapMaterializationHoleRepairs (line 330, loop at line 343) keeps the oldest 24 buckets. The loop at line 570 repairs the ranges oldest first.
  • TimescaleSupport.cs:6109-6117: the probe looks only at [l.mx + 1 hour, s.mn). With no seam row, the slot returns LEAST(l.mn, s.mn).

The seam lies below the successor's floor. So the oldest seam bucket that a repair writes becomes the new s.mn, and the probe window ends there. A seam bucket that is still unrepaired now sits above the new floor. Neither rollup holds it, and the probe does not look at it.

RollupBackfill.cs:39-44 describes the same hazard and runs the backfill newest-first for this reason: "Oldest-first ... the gate armed over a window raw held the only copy of". The seam window reuses the walk's oldest-first order. Oldest-first is safe for interior holes, because an interior repair cannot move the floor. It is not safe for the seam.

Sequence:

  1. On a 3.8.0 store, the query_stats_hourly refresh fails for 3 days while collection continues. The 3.8.0 gate reads only the legacy's floor, so the raw purge stays armed. Raw keeps the rows from the stall because they are less than 4 days old.
  2. Upgrade start (C). The legacy is frozen with l.mx about 3 days back. The successor's first refresh puts s.mn at about C minus 1 day. The seam holds about 48 hourly buckets of rows. The probe finds them, so the gate reads Short and holds the purge. This step is correct.
  3. The next start (C2) is 2 days later. The seam window finds 48 holes in one range. The cap repairs the oldest 24, and s.mn moves down to l.mx + 1 hour.
  4. At the next hourly evaluation, the probe window [l.mx + 1 hour, s.mn) is empty. The slot returns l.mn, so the gate reads Covered and arms the raw purge.
  5. The other 24 buckets are now 3 to 4 days old. If no start repairs them before they pass 4 days, the raw job drops them. If they are inside the 4-day horizon, a later start scans them. Otherwise no start scans them again. After step 3, seamFloor is not below the floor, so the seam window is gone. The ordinary window starts at the horizon (MaterializationHoleScanWindows, lines 400-409).
  6. Loss: about a day of query_stats and procedure_stats history. The frozen legacy stops at l.mx, and the successor does not have these hours. No rollup holds them.

A narrow seam can also cause the same release if it has two or more ranges. Suppose the older range is repaired and then a newer range fails. The refresh can throw, and then the catch at line 624 skips the rest of the target. A shutdown can cancel it. It can also leave buckets, and then remaining > 0 at line 582 logs and continues. In each case the older range has already moved the floor.

Trigger: in the three shapes the brief names, the seam is only the tail of about 2 hours. That tail is one range of at most 3 buckets, so the cap never splits it. The cap trigger needs a refresh that stalls for about 2 days or more while collection continues. The stall can be the legacy's refresh before the upgrade. It can also be the successor's first refresh after the upgrade, for example when its policy failed to attach at the upgrade start. The failure trigger needs a seam with two or more ranges.

Fix: repair the seam window newest-first. Take its cap from the top of the seam (the buckets next to s.mn), and walk its ranges in descending order. Stop the seam walk at the first range that throws or leaves buckets.

Then the floor only grows down without a gap, and every unrepaired seam row stays inside the probe window. This is the property that RollupBackfill states for its slices. Keep oldest-first for the ordinary window. Add a live test: after one walk over a seam of 30 or more hole buckets, IsRawTierDropSafeAsync must still return false.

H2. The stitch trusts the frozen legacy's own span, including holes from before the upgrade that nothing will repair

Where:

  • TimescaleSupport.cs:6109-6117: the probe starts at l.mx + 1 hour. Everything below l.mx is trusted through LEAST(l.mn, s.mn).
  • TimescaleSupport.MaterializationHoles.cs:138 and the note at lines 148-156: the six frozen views are not in MaterializationHoleTargets.
  • TimescaleSupport.MaterializationHoles.cs:525-537: the successor's seam window starts at l.mx plus one bucket.

Sequence:

  1. On a 3.8.0 store, PostgreSQL is down from Monday 10:00 to Wednesday 10:00. The legacy's last bucket is Monday 08:00. Raw holds a tail from Monday 09:00 to 10:00.
  2. After Wednesday, the 1-day policy window never reaches Monday. 3.8.0 has no walk (RepairMaterializationHolesAsync is not on main). The tail becomes a hole inside the legacy's span.
  3. On Thursday at 10:00 the store takes this build after a short stop. l.mx is about Thursday 08:00. The successor's first window starts at about Wednesday 11:00, so s.mn is below l.mx + 1 hour. The seam is empty, the slot returns l.mn, and the gate reads Covered.
  4. On Saturday the raw job drops the Monday chunk. The Monday tail is in neither rollup, so it is gone.

Before 3cfeeea2, the gate named only the successors. It held these rows, and all other raw rows, forever. The stitch now claims coverage for them. In an unfrozen legacy, this build's walk repairs such a tail at the next start, because the tail is inside the 4-day horizon. The freeze takes the legacy out of the walk.

3.8.0 loses the same rows on the same schedule, so this is not a regression against the released build. It meets the brief's High rule because the gate reports Covered over rows that no rollup holds.

Fix (a design decision): use one of these two options.

  • Check the legacy's span bucket by bucket. The span is [time_bucket(source_oldest), l.mx]. An hour with filter-admitted raw rows and no bucket in either rollup makes the slot Short. A matching repair window then refreshes the successor at those hours. Do H1's fix first, and make the seam probe bucket-level too. A repair below l.mx moves s.mn below l.mx, which empties the floor-bounded seam probe while the seam can still hold rows.
  • Accept parity with 3.8.0, and say so in the doc for RetentionArmSafetySql.

M1. An armed raw purge runs when PostgreSQL starts, and it drops a seam older than 4 days before the service can hold or repair it (pre-existing)

Where:

  • TimescaleSupport.cs:5997 (SetRetentionScheduleSql): arm and hold only set the TimescaleDB job's scheduled flag. The job does not check coverage when it runs.
  • Darling/PerformanceMonitor.Darling.Service/DarlingWorker.cs:1959: the service holds the purge again only when its start sweep runs.
  • TimescaleSupport.MaterializationHoles.cs:118-122 calls the ordering against retention "not load-bearing".

Sequence: a 3.8.0 store with an armed raw purge has PostgreSQL stopped for 6 days. When PostgreSQL starts, the TimescaleDB scheduler runs the overdue retention job for query_stats and procedure_stats. The service's start sweep runs later.

drop_chunks removes the chunk that holds the tail of about 2 hours, because that chunk ended more than 4 days ago. After that, the seam window and the probe find nothing. The repair below the horizon that c14b46a0 added has nothing to repair.

These commits do not cause this, and 3.8.0 loses the same tail. So I rank it Medium, although it meets the letter of the High rule. The problem for this PR is a claim. The doc for RetentionArmSafetySql (TimescaleSupport.cs:6055-6058) says that a seam older than the horizon "still gets repaired". That is true only on a store whose raw purge was already held when the store stopped.

Fix: hold the raw retention jobs in the service's graceful shutdown path, and let the start sweep arm them again after it measures. This does not cover a crash or a stop of PostgreSQL alone. A full fix makes each drop run check coverage first, as a custom job. We recommend a separate issue for that, and a sentence in the doc that states the limit.

M2. After an outage of more than a day across the upgrade, the gate holds until the second start

Where:

  • TimescaleSupport.MaterializationHoles.cs:507-512: the walk skips an aggregate with no materialized bucket.
  • Darling/PerformanceMonitor.Darling.Service/DarlingWorker.cs:1940: the walk runs once per start.
  • TimescaleSupport.cs:6059: the doc says the gate "releases automatically, with no manual step".

Sequence: a 3.8.0 store is stopped for 2 days and then takes this build. At the first start, the successors are new and empty, so the walk skips them, and the gate reads Short. After the successor's first refresh, the seam holds the tail of about 2 hours, and each hourly evaluation reads Short. Nothing repairs the tail until the next start.

Until then, the raw purge for both tables stays held, and raw grows by a day of collection each day. No rows are lost. But a store that seldom restarts needs a manual step (a restart or --backfill-rollups) in the exact shape that the brief names.

Fix: run only the seam window, with the same cap, from the hourly retention evaluation when the probe finds a seam. Another option is to run the walk once more after the successor's first materialization. If you choose neither, change the doc to say that the release needs a second start.

L1. Every upgrade from 3.8.0 holds the raw purge again, with a warning that tells the operator to backfill

Where: TimescaleSupport.cs:6110 (COALESCE(s.mn, 'infinity'::timestamp)), and the warning at TimescaleSupport.cs:6676.

Sequence: a healthy 3.8.0 store has an armed purge and no outage. At the upgrade start the successors are empty, so the probe window runs from l.mx + 1 hour to infinity. It always finds the rows after the legacy's last bucket. The gate reads Short, holds the purge, and logs "RE-HELD ... Backfill that consumer".

The successor's first refresh clears it within about 2 hours. No rows are lost. But the warning sends the operator to a heavy verb for a state that clears by itself.

Fix: when s.mn is NULL, end the probe window at now() - HourlyRefreshStartOffset. The successor's first policy window reaches every newer row, and the purge drops only chunks older than 4 days. So those newer rows are not at risk.

L2. A repaired seam older than 3 days never reaches the successor daily, so the successor hourly's 90-day retention is held

Where:

  • TimescaleSupport.MaterializationHoles.cs:525-537: LegacyOf maps only the hourly successors, so a successor daily gets no seam window. The daily's ordinary window starts at its own floor.
  • TimescaleSupport.cs:6213-6215: the retention for each successor hourly waits on its successor daily.

Sequence: after a 5-day outage, the second start repairs the seam into query_stats_interval_hourly. The seam is 5 or more days old. So it is outside the daily's 3-day policy window, and it is below the daily's floor. query_stats_interval_daily never materializes those hours. The hourly's min(bucket) is now older than the daily's, so the hourly's 90-day policy reads Short. It is held, with a warning, until someone runs --backfill-rollups.

No rows are lost, because the hourly keeps them.

Fix: give each successor daily a seam window that starts at its hourly's floor. Or make the seam repair also refresh the dependent daily over the repaired days.

L3. The probe filter for the db-grain slot is narrower than that successor's WHERE

Where: TimescaleSupport.cs:6111 uses IntervalHonestSourceFilter for every slot. query_stats_db_interval_hourly (TimescaleSupport.cs:2147-2148) also requires delta_worker_time IS NOT NULL, and the walk copies that WHERE exactly.

A seam row with a NULL delta and a non-zero interval makes the probe find a row that the walk never treats as a hole. The purge then holds forever. This fails in the safe direction. The current collector writes non-null deltas (CalculateDeltaWithSeriesAge returns long), so only old rows can trigger it.

Fix: build the filter for each slot from MaterializationHoleSourceFilterFor over that successor's CREATE, so the probe and the walk cannot differ.

Answers by question

Q1: yes, see H1, H2 and M1. The zero-interval filter from 3cfeeea2 is safe. All three consumers of the two raw tables exclude those rows. So the filter moves source_oldest later only past rows that no consumer needs. The compression steps drop no rows.

A fresh install is safe. The frozen views stay empty, so l.mx is NULL, the probe's comparison is UNKNOWN, and the slot is the successor's floor alone. A normal field upgrade is safe, with a short hold (L1). For a stop of more than a day across the upgrade, the gate holds until the seam is repaired (M2). After a stop of more than about 5 days, M1 applies.

Q2: yes, see H1.

Q3: nothing new. A seam bucket does not exist before the repair, so compression cannot reach it first. The seam refresh writes below the successor's floor, where no compressed batch exists. compress_after is 2 days for the hourly tier and 4 days for the daily tier (TimescaleSupport.cs:6817 and 6827). The scan horizon is 4 days, so an interior hole that is 2 to 4 days old can sit in a compressed chunk. That case predates this round.

The code has its own measurement on 2.28.1 (TimescaleSupport.cs:6772-6781). A refresh over a compressed range decompresses the overlapping batches and leaves a partial chunk. It does not error or skip. I did not check this on 2.30.x, because there is no rig.

A refresh that stops short does not fail quietly. The walk scans again, forces a refresh, and logs what remains (TimescaleSupport.MaterializationHoles.cs:578-610). In a seam with one range, the floor does not move, so the gate stays held. H1 is the exception.

One point is unverified. The walk's refreshes do not lift timescaledb.max_tuples_decompressed_per_dml_transaction. QueryStoreSliceRepair.cs:322 and Darling/PerformanceMonitor.Darling.Service/DarlingRetention.cs:1242 both lift it. If the internal DML of a refresh counts against that limit, a repair over a compressed chunk on a large store errors. The walk logs that error.

Q4: nothing. The state probe reads only the six frozen names in collect (TimescaleSupport.cs:8065-8090). The drain removes a policy only when a job exists and no chunk is uncompressed (lines 8103-8111). It calls remove_compression_policy with if_exists => true on a frozen name (line 8117). The new guard at line 8867 is a continue, so it can only reduce what the converge changes.

A frozen view keeps a running policy until every chunk is compressed, by design. Nothing refreshes a frozen view, so it gets no new chunk. A paused job never compresses, so the drain never removes it. A missing job leaves nothing to drain. Both of these cost disk only. Compression and policy removal remove no rows.

Q5: nothing for purge or retention. Every caller of For() is a read:

  • ComposeSourceRouter.cs:328 and 330
  • DarlingHealthReader.cs:420
  • DarlingMcpTrendTools.cs:383
  • DarlingTrendReader.cs:931
  • the viewer's DailySummary, FinOps.Workload and QueryTrends files

The purge and retention paths use only RetentionArmSafetySql. They reach it through IsRawTierDropSafeAsync (the catalog sweep at DarlingRetention.cs:306) and EnsureRetentionPoliciesAsync. A note for reads, not a finding: For() takes the deeper floor and does not check the seam. So a read over the seam, or over an H2 hole, routes to the hourly tier and gets no rows while raw still holds them.

22cbb028 and 35a0aefb change only a doc comment and a log message. They have no effect on purge, retention or compression.

@erikdarlingdata

Copy link
Copy Markdown
Owner Author

Round-3 findings below High are filed: M1 as #4299, and M2, L1, L2 and L3 as #4300. H1 and H2 wait on decisions before any fix lands here.

@erikdarlingdata

Copy link
Copy Markdown
Owner Author

H1: a fix lane is running with the ruled fix (the seam walks newest-first and stops at the first failed range) and a live test. H2: ruled parity with 3.8.0. The same lane states that in the RetentionArmSafetySql doc, and #4301 tracks the full bucket-level fix.

…4186 round-3 H1)

A seam wider than one start's repair cap, or with two or more ranges, could
move the successor's floor past a still-unrepaired seam bucket: oldest-first
capping and repair let the OLDEST buckets close first, which drags the
floor (a bare min(bucket)) down to them even while newer seam buckets in
between stay holes, and RetentionArmSafetySql's probe stops looking above
the new floor. The raw purge could then arm and drop rows neither rollup
ever held.

Fix: the seam window's holes are now scanned, capped and walked separately
from the ordinary window's. The seam takes its cap from the newest end
(closest to the successor's floor) and walks descending; the ordinary
window is unchanged (oldest-first, same cap, same continue-past-a-remainder
handling, since an interior repair can never move the floor). The seam
walk stops at the first range that throws (propagates to the existing
per-target catch) or leaves buckets standing (remaining > 0) rather than
moving on to an older range. Both windows share one cap budget per
aggregate per start, seam first, so "a start never re-materializes more
for one aggregate than an ordinary policy run does" stays true.

CapMaterializationHoleRepairs gains a newestFirst parameter (default
false, unchanged behavior) that also splits a straddling range at its
older edge instead of its newer one, so the kept portion stays adjacent
to whatever is already materialized.

Also corrects RetentionArmSafetySql's doc: it claimed an over-horizon seam
"still gets repaired" and the gate "releases automatically, with no manual
step" unconditionally. Both are true only when the raw purge was already
held when the store stopped (#4299) and the outage did not cross the
upgrade (#4300) respectively. Adds the H2-parity note: the stitch trusts
the frozen legacy's own span without checking it, which can read Covered
over a pre-upgrade hole nothing here repairs -- accepted because 3.8.0
loses the same rows on the same schedule, and the A6 backfill
(--backfill-rollups) covers it for an operator who runs it.

Two live tests against a real TimescaleDB rig:
- a 35-bucket seam (over the 24-bucket cap) stays Short after one walk
  (top 24 repaired, bottom 11 still a hole) and reads Covered only after
  a second walk closes the rest; revert-proved by temporarily restoring
  oldest-first, which fails at the same assertion.
- a two-range seam where the newer range's refresh is forced to raise
  23514 via a CHECK constraint on its own materialization chunk: the
  older range is never touched and the gate still reads not-safe.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
@erikdarlingdata

erikdarlingdata commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner Author

Round-3 H1 fix, pushed to fix/3653-a6-freeze at c3291f6

Commit: c3291f66 (parent 35a0aef, the branch head named in the review). No PR opened. None of #4186's title, body or state touched. #3653 untouched. CHANGELOG.md untouched.

1. Seam repairs newest-first, stops at the first failed range

Darling/PerformanceMonitor.Darling.Storage/TimescaleSupport.MaterializationHoles.cs:

  • CapMaterializationHoleRepairs (line 338) gains a newestFirst parameter, default false (existing callers and behavior unchanged). When true it orders descending and splits a straddling range at its OLDER edge, so the kept portion stays adjacent to whatever is already materialized.
  • The per-target loop (from line ~566) now scans the seam window's holes separately from the ordinary window's. isSeamWindow (line 583) is true exactly when seamFloor < floor.Value. That is also exactly when MaterializationHoleScanWindows emitted a seam window as windows[0].
  • The seam is capped with newestFirst: true at line 616. The ordinary window keeps the old call, unchanged.
  • One decision the ruling left open: whether the seam gets its own cap budget or shares today's one. I read "keep today's cap... for the ordinary window" as keeping the total pool the same size. So both windows share one cap per aggregate per start, seam spent first. That keeps the ordinary window's existing log line ("a start never re-materializes more for one aggregate than an ordinary policy run does") literally true. Flagging this in case the intent was two independent budgets instead.
  • The seam walk (line 684) breaks after any range that leaves buckets standing (remaining > 0). A thrown exception, from either window, already propagated to the per-target catch before this change. So newest-first ordering alone makes that arm safe too, with no separate catch needed inside the seam loop.

2. Gate doc (RetentionArmSafetySql, TimescaleSupport.cs)

3. Tests

  • MaterializationHoleRepairTests.cs:213 TheCap_NewestFirst_...: unit-pins the newest-first split arithmetic (mirrors the existing oldest-first pin at line ~177 with the same three ranges).
  • FrozenRollupLiveTests.cs:648 Outage_SeamWiderThanTheCap_...: live, rig-run. Seeds a 35-bucket seam (cap is 24). One walk repairs the top 24 and leaves 11 below the new floor. IsRawTierDropSafeAsync is asserted still false. A second walk closes the rest and it is then asserted true. Revert-proved: with newestFirst: false put back on the seam call, this test fails at the first IsRawTierDropSafeAsync assertion, because the gate reads true (Covered) after only one walk. I confirmed this by rebuilding and running it, then restored the fix and re-confirmed green.
  • FrozenRollupLiveTests.cs:739 Outage_SeamTwoRanges_...: live. Builds a seam with two distinct ranges (older A, newer B). A CHECK constraint added directly to B's own materialization chunk forces B's refresh to raise error 23514. This is the "throws" arm the brief names as an acceptable alternative to "leaves buckets". I verified the constraint mechanism empirically on the rig first: ALTER TABLE against the aggregate's parent hypertable is refused, but a concrete chunk table takes the constraint. The test asserts summary.Failures >= 1, that A has zero rows in the successor (never attempted), and that the gate still reads false.
  • Existing classes kept green: MaterializationHoleRepairTests (18 tests), FrozenRollupLiveTests (10, all), IntervalHonestHourlyRollupTests + TimescaleContinuousAggregateTests + RollupBackfillTests (83 combined), PayloadDimensionLiveTests + QueryStoreCorrectedRollupLiveTests (25 combined), DocCommentHygiene (77). One pre-existing pin (Worker_WritesTheSummaryLine_...) needed a deliberate update, since it string-matches the holesFound += / bucketsFound += source lines verbatim. I updated it to the new, equivalent (seam plus ordinary, summed) text and noted the reason inline.
  • Build: Darling.Tests.csproj -c Debug, 0 Warning(s), 0 Error(s), each time.

Not run: the full suite

Context ran past the 150k mark the brief sets for a full-suite run, partway through building the failure-arm test's constraint mechanism. Everything above ran as targeted classes only, all green, against a fresh darlingtest on the rig (port 55999, stopped with pg_ctl -m fast stop when done). The coordinator-side PR tender should run the full suite once before treating this as final. Installer.Tests was not run, per the rule.

@erikdarlingdata erikdarlingdata changed the title DO NOT MERGE (A6 freeze, lane LC): legacy trio off the refresh grid Legacy trio off the refresh grid (#3653) Sep 26, 2026
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review September 26, 2026 05:51
@erikdarlingdata
erikdarlingdata merged commit 904d460 into dev Sep 26, 2026
17 of 18 checks passed
@erikdarlingdata
erikdarlingdata deleted the fix/3653-a6-freeze branch September 26, 2026 05:51
erikdarlingdata added a commit that referenced this pull request Sep 26, 2026
…floor

The walk's #4186 seam-fix block lowered seamFloor to the frozen legacy's
last bucket + one bucket width. Per the ruling (#4301 comment 5844202704),
the walk now lowers it to raw's own filtered floor (AlignDown'd), the same
bound RetentionArmSafetySql's gate probes from, whenever that reaches
further back than the successor's own floor. Everything downstream (the
successor-only scan, the newest-first cap and walk, the shared cap with
the ordinary window) is unchanged and is itself the contiguous downward
fill: RollupCoverage.StitchedRelationSql splits its read at the
successor's floor, so every row below it must land as a successor bucket
for the stitch to read each row exactly once.

Removed the now-dead legacyInteriorFrom/legacyInteriorTo third-window
parameters from MaterializationHoleScanWindows and the two pins that
exercised them (ScanWindows_LegacyInterior_*, ScanWindows_NoLegacyInterior_*):
under this ruling a legacy-interior hole is repaired by the same
newest-first seam descent as everything else below the successor's floor,
so the separate oldest-first interior branch never gets built. Removed
LegacySuccessorHoleScanSql (the walk's list-form twin of the gate's
LegacySuccessorHoleExistsSql) since the walk no longer shares that hole
definition; RetentionArmSafetySql keeps LegacySuccessorHoleExistsSql for
its own probe.
erikdarlingdata added a commit that referenced this pull request Sep 26, 2026
…the successor down to raw's floor (#4401)

Closes a gap in the raw-purge gate after the legacy rollups were frozen (#4186).

- RetentionArmSafetySql no longer trusts the frozen legacy rollup's whole span. It checks for a hole inside that span, not only for a seam above it, before reporting the raw purge covered.
- The repair walk fills the successor rollup contiguously downward to raw's own floor, newest first, instead of stopping at the legacy's last materialized bucket. An isolated repair below the legacy's last bucket would move the successor's first bucket down, so stitched reads would take the hours in between from the successor, which doesn't hold them.
- FrozenRollupLiveTests has five new pins. Three existing seam tests have their expected counts and floors updated for the contiguous fill; no assertion was loosened.

Refs #4301
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant