Skip to content

After an outage across an upgrade, the raw purge gate releases within the hour on a running store (#4300) - #4430

Merged
erikdarlingdata merged 5 commits into
devfrom
fix/4300-seam-repair-on-hourly-tick
Sep 26, 2026
Merged

erikdarlingdata merged 5 commits into
devfrom
fix/4300-seam-repair-on-hourly-tick

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Refs #4300 (item 1).

Why

The raw-tier purge gate stays held until every hole between a frozen legacy rollup and its successor is repaired. Before this change, the only thing that ever re-ran that repair walk was a full service start — a legacy/successor seam an outage opened across an upgrade would sit un-repaired, and the gate would stay Short, until the store was restarted a second time. A store that keeps running after the upgrade could hold the raw purge indefinitely with no restart in sight.

Why the earlier relaunch doesn't cover this

A prior, related fix (#4391) changed when the periodic repair relaunches the full hole-scan walk, but only when the store's repair-epoch stamp is stale under the current postmaster start. A first start whose repair completed cleanly (0 failures, the successor skipped because it had materialized nothing yet) stamps that epoch, so the periodic pass never relaunches the full walk — #4391 does not cover this seam shape.

What changes

  • TimescaleSupport.MaterializationHoles.cs: the per-target hole-repair walk is factored into a shared RepairMaterializationTargetsAsync, with a new RepairMaterializationSeamsAsync — a seam-only variant restricted to legacy-paired targets and their seam window, using the exact same per-target body, cap and failure isolation as the full walk. RepairMaterializationHolesAsync delegates to the same shared method unchanged.
  • DarlingWorker.ReevaluateRetentionPoliciesAsync (the hourly Periodic pass) now runs the seam-only repair before the coverage sweep re-judges the gate, so a seam closed this tick is already gone when the gate is re-evaluated a moment later.
    • Its own budget: the repair runs under a 2-minute child budget linked to the pass's 5-minute budget. If the repair uses its 2 minutes, it stops and logs one Warning naming the hours still deferred. The pass then continues to the coverage sweep, the purge trigger and the epoch relaunch, so a slow store is never starved of them. A cut-short refresh resumes safely on the next pass. If the deferred-hour count itself can't be read, the Warning says "unknown" and the pass still continues.
    • No overlap: while the start-path repair is still running, the hourly seam repair skips the pass. The full walk covers the seam.
    • Failure-isolated: any other throw logs a Warning, and the pass continues. A shutdown still propagates to the pass's outer catches.
    • The seam pass never stamps the repair epoch. Only the full start-path repair does.
  • The storage doc that used to describe this as needing a second start now describes the hourly seam repair closing the gate within the hour, with no restart, citing #4186 gate follow-ups: release needs a second start after an outage, plus three smaller holds #4300.

Test plan

  • The single-start pin's RED on dev: a throwaway probe test, copied from the branch's new Outage_SeamBetweenFrozenLegacyAndSuccessor_SingleStart_HourlySeamRepairReleasesTheGate and adapted to what dev's hourly pass actually does (only EnsureRetentionPoliciesAsync, no repair of any kind — RepairMaterializationSeamsAsync does not exist on dev), built and run in a detached worktree of origin/dev: after one start's repair skips the still-empty successor, the successor's first refresh runs on the running service, and a dev-shaped hourly tick (coverage re-judged, no repair in between) is simulated. The assertion that the gate should read Covered fails with on dev, the single-start hourly tick never runs a seam repair, so the gate stays Short — confirmed RED.
  • GREEN on the branch: the real test Outage_SeamBetweenFrozenLegacyAndSuccessor_SingleStart_HourlySeamRepairReleasesTheGate passes: one start, no restart, the hourly seam-only repair alone closes 7 buckets and the gate reads Covered.
  • The seam-only fixture fix: SeamOnlyRepair_NeverTouchesAnOrdinaryInteriorHoleOnANonLegacyTarget was failing on its own trailing sanity check (the full walk should have closed the interior hole the seam-only call left standing) because the fixture's gap had no source row in it — a bucket with nothing in the source is legitimately empty, not a hole, under this repair's own definition of a hole. Seeded a real row in the gap bucket instead of weakening the assertion; now the full walk genuinely closes it and the test passes.
  • The doc pin: a new fact in RetentionReevaluationTests reads the storage source (the repo's existing ReadStorageSource() helper) and asserts it no longer contains "SECOND start".
  • The worker-shape pin (Worker_ReevaluationMethod_RunsThePeriodicPass_UnderOneBudget_FailureIsolated) now:
    • finds the pass's OUTER shutdown catch (the last one before the budget catch), past the seam repair's inner one;
    • checks the order shutdown → budget → everything else, and no rethrow from the outer shutdown catch onward;
    • checks that the inner shutdown catch rethrows;
    • checks that the seam call comes before the coverage sweep, and that no full hole-repair walk is called on the hourly pass;
    • pins each Warning's text.
    • Mutation: adding throw; to the outer shutdown catch fails it (Assert.DoesNotContain() Failure: Sub-string found … "throw"). Restored, it passes.
  • The single-start pin also asserts HolesRemaining, BucketsDeferred and Failures are all 0.
  • A seam wider than the per-pass cap (a new live pin): the seam-only repair closes 24 buckets and defers 12. The gate stays Short, and no repair epoch is stamped.
  • Ran live, on a local TimescaleDB 2.30.1-pg18 rig (latest head): FrozenRollupLiveTests 24/24, RawPurgeTriggerLiveTests + RetentionReevaluationLiveTests 9/9, RetentionReevaluationTests 10/10, MaterializationHoleRepairTests 20/20, DocCommentHygieneTests 77/77.
  • Both Darling.Tests.csproj and Lite.Tests.csproj build 0 errors with -p:EnableWindowsTargeting=true; Darling.Tests/Lite.Tests target net10.0-windows and cannot execute on macOS outside the in-process runner used above — CI decides the Windows-only suites.

CHANGELOG entry

SECTION: Fixed
ENTRY:

… a doc pin (#4300)

- The interior-hole fixture in SeamOnlyRepair_NeverTouchesAnOrdinaryInteriorHoleOnANonLegacyTarget
  left a bucket with no source row between two materialized spans, which is legitimately empty, not
  a hole -- the full walk's sanity assertion failed because there was nothing to repair. Seed a real
  row in the hole bucket so the full walk actually closes it.
- RetentionReevaluationTests: the seam repair's own failure-isolated try/catch (added ahead of the
  coverage sweep) sits before the outer shutdown/budget/everything-else catches the existing pin
  checked the order of, and adds a third LogWarning site. Scoped the ordering/no-rethrow check to the
  outer catches and updated the warning count and text to match.
- Added a doc pin (RetentionReevaluationTests) asserting the storage source no longer says the gate
  needs a SECOND start.
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review September 26, 2026 16:06
@erikdarlingdata
erikdarlingdata merged commit dc2027b into dev Sep 26, 2026
17 of 18 checks passed
@erikdarlingdata
erikdarlingdata deleted the fix/4300-seam-repair-on-hourly-tick branch September 26, 2026 16:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant