Skip to content

Dimension GC can prune while a FAILED fact purge leaves references alive #1782

Description

@erikdarlingdata

Derived from a field question that carried a false premise worth correcting on the record first: the payload dimension tables DO have a GC (product-side, not a TimescaleDB policy - it will never appear in the policy catalogs). DarlingRetention.cs runs every dim table through the time-sliced DELETE on last_seen (call site :226), with cutoff = now - (widest fact retention + ChunkIntervalDays + 1) resolved through the same retentionDaysFor resolver the fact purge uses (:210-:225). For the dim-feeding tables that resolves to ~32 days today. The upsert refreshes last_seen (at most hourly) for any digest still being collected, so still-live content never ages toward the cutoff. Shipped in #1768, adversarially reviewed, gated-live tested (Gc_DeletesDimRowsPastTheHorizon_AndKeepsTheOnesInsideIt).

The real gap the question surfaced

The cutoff arithmetic embeds an assumption: facts older than their retention are GONE. That is true when the purges run - and FALSE on any store where the tiered raw purges are HELD by the #1680 coverage gate (the #1759 state, which is exactly the live state of a production field instance today). On such a store, digest-carrying fact rows can outlive the policy window indefinitely. A digest whose query stopped running 32+ days ago, but whose held fact rows still reference it, gets its dimension row pruned while live references remain: the resolving view returns NULL payload for those rows, and the D2b unresolvable-digest check goes nonzero.

Fuse math on the field box: digest-carrying rows began 2026-07-27, so the earliest possible bad prune is ~32 days out (~2026-08-28), and only if the raw purges are STILL held then. #1759 Phase 2 (in flight) arms them well before that on any box that runs the backfill. So: LOW severity, real coupling defect.

Fix shape (small, fail-safe)

The held state is knowable from the product's own arming-gate verdicts. While any dim-feeding table's raw purge is HELD: either SKIP the dimension GC for that cycle with one log line ("dimension GC deferred: raw purges held, facts may outlive their policy window"), or clamp the cutoff to (oldest digest-carrying collection_time - margin). Skip-with-log is simpler and honest; the clamp needs a bounded query to be worth it. Either way, add the mutation: simulate the held state with old digest-carrying facts present, run the GC, and the dim row they reference MUST survive - watched red against the current cutoff.

Also confirmed by the same field pass (the good half of the premise)

The continuous aggregates carry ZERO payload/digest columns (verified via information_schema on all hourly/daily/baseline aggregates) - the pre-dedup bloat is strictly confined to raw chunks, which age out on the existing 4-day schedule with no manual work. The damage boundary is as designed.

Activity

  1. erikdarlingdata commented on Jul 28, 2026

    @erikdarlingdata
    OwnerAuthor

    Retitled: the original premise does not hold, but a real trigger does. Full derivation, verified against origin/dev @ 50012266.

    Why a HELD policy cannot cause this

    There are two independent purges on these tables:

    1. The tiered policy — RawRetentionInterval = "4 days" (TimescaleSupport.cs:1223), created paused, armed by the Retention policies run their first check immediately — no external window to pause them, caused permanent data loss #1680 coverage gate. This is the one that is HELD.
    2. Darling's own catalog sweep — DarlingRetention.PurgeAsync iterates CollectorCatalog.All and drops chunks at the per-collector horizon (DarlingRetention.cs:143-175). query_stats and procedure_stats are both new(1, 30) (CollectorScheduleDefaults.cs:42-43) → 30 days. This loop has no coverage gate and no held check; the only conditional is per-table failure isolation.

    The dimension GC cutoff is widest(30) + ChunkIntervalDays(1) + 1 = 32 days (DarlingRetention.cs:247).

    So a held policy changes whether facts live 4 days or 30 — both inside the 32-day cutoff. Facts are removed at 30 by a purge that is not held; dimensions are pruned at 32. Dimensions outlive their facts by two days by construction, and the ordering survives any retention override because both sides resolve through the same retentionDaysFor resolver (raise query_stats to 90 and the cutoff becomes 92).

    The original fuse math ("earliest bad prune ~2026-08-28") assumed facts persist past 32 days, which requires the 30-day sweep to also not remove them.

    The trigger that IS reachable

    The fact loop runs BEFORE the dimension GC in the same method, and is failure-isolated: a table whose drop_chunks fails, and whose DELETE fallback then also fails, is warned and skipped while the sweep continues into the dimension GC. Facts persist past 32 days; the GC prunes on schedule; digests dangle and the resolving view serves NULL payload with no error.

    So the coupling defect is real — it is keyed to a purge FAILURE, not to a held policy.

    Fix

    Defer the dimension GC for a cycle when a dim-feeding table's fact purge failed in that same cycle. The sweep already knows which tables failed, so it costs no extra query, defers only when something actually went wrong, and self-ends on the next successful sweep — where a held-policy check would defer indefinitely on a store whose policies are held for unrelated reasons.

    Spun off

    Establishing the above surfaced a larger, opposite-signed defect in the same file: the ungated catalog sweep bypasses the #1680 coverage gate and destroys rollup-uncovered history at 30 days, on a dated fuse. Filed as #1784.

  2. changed the title [-]Dimension GC cutoff assumes fact purges run on schedule; a HELD tiered purge voids the assumption[/-] [+]Dimension GC can prune while a FAILED fact purge leaves references alive[/+] on Jul 28, 2026
  3. erikdarlingdata commented on Jul 28, 2026

    @erikdarlingdata
    OwnerAuthor

    Fixed by #1789 (merged to dev at c234a3a): the payload-dimension GC now defers when a dim-feeding purge failed, so a failed fact purge can no longer leave references alive for the GC to prune out from under. The interaction with coverage-gated skips is tracked separately in #1795.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions