Repository navigation
Dimension GC can prune while a FAILED fact purge leaves references alive #1782
Description
Activity
Retitled: the original premise does not hold, but a real trigger does. Full derivation, verified against
origin/dev@50012266.Why a HELD policy cannot cause this
There are two independent purges on these tables:
- The tiered policy —
RawRetentionInterval = "4 days"(TimescaleSupport.cs:1223), created paused, armed by the Retention policies run their first check immediately — no external window to pause them, caused permanent data loss #1680 coverage gate. This is the one that is HELD. - Darling's own catalog sweep —
DarlingRetention.PurgeAsynciteratesCollectorCatalog.Alland drops chunks at the per-collector horizon (DarlingRetention.cs:143-175).query_statsandprocedure_statsare bothnew(1, 30)(CollectorScheduleDefaults.cs:42-43) → 30 days. This loop has no coverage gate and no held check; the only conditional is per-table failure isolation.
The dimension GC cutoff is
widest(30) + ChunkIntervalDays(1) + 1 = 32days (DarlingRetention.cs:247).So a held policy changes whether facts live 4 days or 30 — both inside the 32-day cutoff. Facts are removed at 30 by a purge that is not held; dimensions are pruned at 32. Dimensions outlive their facts by two days by construction, and the ordering survives any retention override because both sides resolve through the same
retentionDaysForresolver (raisequery_statsto 90 and the cutoff becomes 92).The original fuse math ("earliest bad prune ~2026-08-28") assumed facts persist past 32 days, which requires the 30-day sweep to also not remove them.
The trigger that IS reachable
The fact loop runs BEFORE the dimension GC in the same method, and is failure-isolated: a table whose
drop_chunksfails, and whose DELETE fallback then also fails, is warned and skipped while the sweep continues into the dimension GC. Facts persist past 32 days; the GC prunes on schedule; digests dangle and the resolving view serves NULL payload with no error.So the coupling defect is real — it is keyed to a purge FAILURE, not to a held policy.
Fix
Defer the dimension GC for a cycle when a dim-feeding table's fact purge failed in that same cycle. The sweep already knows which tables failed, so it costs no extra query, defers only when something actually went wrong, and self-ends on the next successful sweep — where a held-policy check would defer indefinitely on a store whose policies are held for unrelated reasons.
Spun off
Establishing the above surfaced a larger, opposite-signed defect in the same file: the ungated catalog sweep bypasses the #1680 coverage gate and destroys rollup-uncovered history at 30 days, on a dated fuse. Filed as #1784.
- The tiered policy —
- changed the title
[-]Dimension GC cutoff assumes fact purges run on schedule; a HELD tiered purge voids the assumption[/-][+]Dimension GC can prune while a FAILED fact purge leaves references alive[/+]on Jul 28, 2026 - added 4 commits that reference this issue
on Jul 28, 2026
Derived from a field question that carried a false premise worth correcting on the record first: the payload dimension tables DO have a GC (product-side, not a TimescaleDB policy - it will never appear in the policy catalogs). DarlingRetention.cs runs every dim table through the time-sliced DELETE on last_seen (call site
:226), with cutoff = now - (widest fact retention + ChunkIntervalDays + 1) resolved through the same retentionDaysFor resolver the fact purge uses (:210-:225). For the dim-feeding tables that resolves to ~32 days today. The upsert refreshes last_seen (at most hourly) for any digest still being collected, so still-live content never ages toward the cutoff. Shipped in #1768, adversarially reviewed, gated-live tested (Gc_DeletesDimRowsPastTheHorizon_AndKeepsTheOnesInsideIt).The real gap the question surfaced
The cutoff arithmetic embeds an assumption: facts older than their retention are GONE. That is true when the purges run - and FALSE on any store where the tiered raw purges are HELD by the #1680 coverage gate (the #1759 state, which is exactly the live state of a production field instance today). On such a store, digest-carrying fact rows can outlive the policy window indefinitely. A digest whose query stopped running 32+ days ago, but whose held fact rows still reference it, gets its dimension row pruned while live references remain: the resolving view returns NULL payload for those rows, and the D2b unresolvable-digest check goes nonzero.
Fuse math on the field box: digest-carrying rows began 2026-07-27, so the earliest possible bad prune is ~32 days out (~2026-08-28), and only if the raw purges are STILL held then. #1759 Phase 2 (in flight) arms them well before that on any box that runs the backfill. So: LOW severity, real coupling defect.
Fix shape (small, fail-safe)
The held state is knowable from the product's own arming-gate verdicts. While any dim-feeding table's raw purge is HELD: either SKIP the dimension GC for that cycle with one log line ("dimension GC deferred: raw purges held, facts may outlive their policy window"), or clamp the cutoff to (oldest digest-carrying collection_time - margin). Skip-with-log is simpler and honest; the clamp needs a bounded query to be worth it. Either way, add the mutation: simulate the held state with old digest-carrying facts present, run the GC, and the dim row they reference MUST survive - watched red against the current cutoff.
Also confirmed by the same field pass (the good half of the premise)
The continuous aggregates carry ZERO payload/digest columns (verified via information_schema on all hourly/daily/baseline aggregates) - the pre-dedup bloat is strictly confined to raw chunks, which age out on the existing 4-day schedule with no manual work. The damage boundary is as designed.