Repository navigation
Raw purge gate: check the frozen legacy's span bucket by bucket, and repair holes inside it #4301
Description
Activity
Claude posting for Erik Darling
Ruled (2026-09-26): how the raw purge gate treats a hole in the successor rollups that can't be repaired.
The gate separates a hole whose source state is UNKNOWN from one that is PROVABLY unrepairable. It does not hold every unrepairable hole the same way.
- Unknown source state (it can't be determined whether raw rows for the hole's bucket still exist: a probe error, a missing catalog object, an unreadable chunk floor): the purge is held. An unknown answer is not a safe answer, the same rule
IsRawTierDropSafeAsyncalready applies. - Provably unrepairable (the hole's bucket lies entirely below raw's current chunk floor, so its source rows are already gone): the purge is not held. Holding newer raw rows can't recover an old bucket whose source is gone; it would only let raw grow without bound. Instead:
- the hole is recorded as a permanent gap;
- the existing Retention Held-family alert is raised once, naming the table, the bucket range and "source already purged";
- coverage proceeds past the gap.
- A repairable hole is repaired before the purge proceeds, as the issue describes.
Pins:
- (a) repairable hole: the gate reads Short before repair and Covered after;
- (b) unknown source state: the purge is held;
- (c) provably unrepairable: the purge isn't held, the alert is raised once, and the gap is recorded;
- (d) a second pass doesn't raise the alert again.
Order: this issue's lanes run before #4300's, because both change the same probe and scan-window code.
- Unknown source state (it can't be determined whether raw rows for the hole's bucket still exist: a probe error, a missing catalog object, an unreadable chunk floor): the purge is held. An unknown answer is not a safe answer, the same rule
Claude posting for Erik Darling
Ruling amended (2026-09-26): no gap record and no alert for a hole that can't be repaired. This replaces the "provably unrepairable" part of issuecomment-5843156201.
- Why: below raw's floor, a bucket lost to an earlier purge and an hour with no collection leave the same evidence: zero admitted rows. Nothing in the store can tell them apart without a new table.
- Consequences:
- The gate never holds on such a bucket, because it is never a hole. That is already the ruled purge behaviour.
- After this fix, the purge never passes a hole the gate can detect, so no new gaps of this kind can arise.
- No gap record, no alert, and no migration.
- Release text, stated once: gaps left by earlier versions' purges below the raw floor cannot be detected or repaired. From this version, the purge holds until every detectable hole is repaired.
- Pins:
- (a) repairable hole: Short before repair, Covered after;
- (b) unknown source state: the purge holds;
- the seam trap: after an interior repair moves the successor's floor below the legacy's last bucket, a seam hole still reads Short;
- a bucket below raw's floor is never a hole and never holds the purge.
- Unchanged: the gate and the repair walk share one hole definition, per hour bucket, both interior and seam.
- added 5 commits that reference this issue
on Sep 26, 2026 Claude posting for Erik Darling
Ruled (2026-09-26): how the repair walk fills a hole below the legacy's last bucket. It fills the successor contiguously downward, not at isolated hours.
- Why: stitched reads split at the successor's own first bucket. They read the frozen legacy below it and the successor from it upward, so they assume the successor holds every hour from its first bucket up. Refreshing the successor at one isolated hour below the legacy's last bucket moves that first bucket down to it. Reads for every hour between would then come from the successor, which doesn't hold them, and those rows would silently drop out of every stitched read. Today's seam window has the same dependency: it is bounded by the successor's first bucket, so it would disappear after one such repair.
- The repair: for a successor with a frozen legacy, the walk fills the successor over
[time_bucket(raw's filtered floor), successor's first bucket), newest-first from the successor's first bucket downward, under the existing per-pass cap. The successor stays contiguous from its first bucket up, so each stitched row is still read exactly once, and the seam is covered by the same walk.- This replaces the body's "a matching repair window refreshes the successor at those hours".
- The gate is unchanged. It still reads the legacy's interior and the seam bucket by bucket, and holds the purge while any hole is detectable.
- Pins:
- a repairable hole reads Short, the walk repairs it, and it then reads Covered;
- an unknown source state keeps the purge held;
- after a repair, a stitched read over the hours between the repaired bucket and the legacy's last bucket returns every legacy row, identical to before;
- on a store whose successor already reaches below raw's oldest row, the walk refreshes nothing;
- the per-pass cap: the descent resumes on the next pass with no gap.
- added 4 commits that reference this issue
on Sep 26, 2026 Claude posting for Erik Darling.
#4401 is merged. The raw-purge gate now checks the frozen legacy rollup's own span for holes. The repair walk fills the successor contiguously down to raw's floor, so stitched reads stay gap-free.
Claude posting for Erik Darling.
Closing: the work is on dev in #4401 (018cf1f). The raw purge gate checks the frozen legacy rollup's span bucket by bucket, and holds the purge when a hole's source state can't be determined. The repair walk fills the successor contiguously down to raw's floor. No work remains on this issue.
Follow-up to #4186, from its round-3 review (H2). For #4186 the ruling is to accept parity with 3.8.0 and to say so in the doc for
RetentionArmSafetySql. This issue is the real fix, and it comes after #4186's H1 fix.What happens
The raw purge's gate trusts the frozen legacy rollup's whole span. Its probe starts at the legacy's last bucket plus 1 hour (
TimescaleSupport.csnear lines 6109-6117). Everything below that counts as covered, throughLEAST(l.mn, s.mn). The six frozen legacy views are not inMaterializationHoleTargets, so the hole walk never repairs them. SeeTimescaleSupport.MaterializationHoles.csnear line 138, and the note near lines 148-156.3.8.0 loses the same rows on the same schedule. But the gate now reports Covered over rows that no rollup holds.
Fix
Check the legacy's span bucket by bucket, over
[time_bucket(source_oldest), l.mx]. An hour that has raw rows the filter admits, and no bucket in either rollup, makes the slot Short. A matching repair window then refreshes the successor at those hours.Make the seam probe bucket-level too. A repair below the legacy's last bucket moves the successor's floor below it. That empties the floor-bounded seam probe while the seam can still hold rows.