Skip to content

Bound the dimension GC by the oldest surviving digest-carrying fact (#1795) - #1813

Merged
erikdarlingdata merged 2 commits into
devfrom
feature/1795-gc-measured-bound
Jul 28, 2026
Merged

erikdarlingdata merged 2 commits into
devfrom
feature/1795-gc-measured-bound

Conversation

@erikdarlingdata

Copy link
Copy Markdown
Owner

Closes #1795, to the issue's own design notes.

The change

The #1782 guard deferred the WHOLE dimension GC whenever a dim-feeding purge failed — and the #1784 coverage clamp holds those purges every sweep on a coverage-lagging store (the production field state), so the GC deferred every sweep until a backfill landed. A 400-day orphan survived with nothing failed anywhere.

The GC now measures the true safety boundary: content older than the oldest SURVIVING digest-carrying fact row cannot be referenced by anything, whatever the reason those facts survive. Per the issue's notes:

  • Bounded probe: min(collection_time) per dim-feeding table under exactly the predicate of a new V39 partial index per table — an index-edge read once per sweep, not an oldest-chunk walk through pre-query_text/query_plan_xml stored inline per row: 94% of a field store — normalize into hash-keyed dimension tables (~135x measured) #1767 NULL-digest rows. Probe predicate, index predicate, and PayloadDimensions.All are pinned to each other.
  • Minimum across tables, same reasoning as the widest-retention horizon.
  • Margins preserved: the measured side carries the same one-day margin as the assumed side (the hourly last_seen refresh guard); ComputeDimensionCutoff is pure and unit-tested (healthy floor → assumed horizon rules unchanged; held floor → clamped; no digest facts → assumed).
  • The guard becomes unnecessary rather than dormant — it survives in one honest form: an UNMEASURABLE floor (table missing/unreachable) still defers the cycle, because pruning on an unknown boundary is the one way to dangle digests. New fixed log lines for both states (deferred-unmeasurable, bounded-by-surviving-facts).

Viewer: schema ladder gains the V39 arm (index-existence sentinel, the V22 pg_indexes idiom) so a fully-migrated store maps to exactly RequiredStoreSchemaVersion. StorageVersion 38 → 39; all version pins updated.

Fixture defect found and fixed en route

EnsureContinuousAggregatesAsync attaches refresh policies whose jobs fire IMMEDIATELY (#1788's finding). The live class's deep force-refresh collided with them (55P03), and a restore's DROP could collide with a running job's lock — silently stranding query_stats_db_hourly/db_daily in the shared fixture, whose leftover deep coverage then flipped the #1784 gate for every later test (the same manufactured-flake mechanism #1794 documented). The tests refresh manually and never needed the scheduler: the ensure wrapper now removes every rollup's refresh policy immediately, and the one force-refresh that can still catch an already-executing job retries bounded on 55P03 only — the product's own #1788 idiom, not a retry-wrap of an assertion.

Test plan

  • New live test: on a coverage-clamped store the orphan is PRUNED and the referenced content SURVIVES, with the clamp verifiably holding and the bounded signature logged
  • Rewritten deferral test: an unmeasurable floor (renamed table) defers with the new signature; measurable again → prunes
  • Both watched RED by mutating ComputeDimensionCutoff to ignore the measured floor (unit pin + live test both fired)
  • 3x consecutive full-class live runs against PostgreSQL 18.4 / TimescaleDB 2.28.1 — green, ZERO stranded aggregates after
  • Full fast suite 3569 green (all version-ladder pins updated); service + viewer builds, 0 warnings
  • darling-pg CI leg (fresh cluster, runs the V39 migration end to end)

🤖 Generated with Claude Code

erikdarlingdata and others added 2 commits July 28, 2026 14:32
…1795)

The #1782 guard deferred the whole dimension GC whenever a dim-feeding
purge failed - and the #1784 coverage clamp holds those purges EVERY
sweep on a coverage-lagging store, so the GC deferred every sweep and a
400-day orphan survived with nothing failed anywhere.

The GC now measures the true safety boundary instead of assuming one:
min(collection_time) per dim-feeding table under exactly the predicate
of a new V39 partial index (index-edge probe, once per sweep), minimum
across tables, clamping the assumed cutoff to one day before it. Held
history bounds the GC instead of stopping it; referenced content
survives; orphans reclaim. An UNMEASURABLE floor (table missing) still
defers - pruning on an unknown boundary is how digests dangle.

Viewer schema ladder gains the V39 arm (index-existence sentinel).
Probe predicate, index predicate, and dimension map pinned three ways.

Test-fixture defect fixed en route: EnsureContinuousAggregates attaches
refresh policies whose jobs fire immediately; the class's force-refresh
collided (55P03) and a restore DROP could strand db-grain aggregates in
the shared fixture, flipping the #1784 gate for later tests. The ensure
wrapper now removes rollup policies (tests refresh manually) and the
force-refresh retries bounded on 55P03 only - the product's #1788 idiom.

Verified: new live test proves orphan-pruned + referenced-kept while the
clamp holds; rewritten deferral test proves the unmeasurable-floor path;
both watched RED by mutating the cutoff to ignore the floor; 3x
consecutive full-class live runs, zero stranded aggregates; full fast
suite 3569 green; service + viewer builds zero warnings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The darling-pg leg (fresh cluster) failed the deferral test with
"Connection is not open" thrown from cleanup - the body succeeded, a
RestoreCaggs drop died mid-cleanup, TryExecAsync swallowed it silently,
and the broken session took out the next helper. Exactly the #1810
contract working (the real failure surfaced loudly) - the defect was in
this class's best-effort helper.

- TryExecAsync reopens the connection if a prior swallowed failure
  closed it (fresh pooled session), and prints what it swallowed -
  xUnit captures console per test, so the next mystery is a diagnosis.
- The deferral test's cleanup runs its dim delete BEFORE the CAGG
  restore: the drops are the one step that can break the session, and
  a break there now has no cleanup statements left to strand.

Verified: 2x full live class green on the local rig.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@erikdarlingdata
erikdarlingdata merged commit ff3f778 into dev Jul 28, 2026
4 checks passed
@erikdarlingdata
erikdarlingdata deleted the feature/1795-gc-measured-bound branch July 28, 2026 18:48
pull Bot pushed a commit to ehtick/PerformanceMonitor that referenced this pull request Jul 29, 2026
…ikdarlingdata#1815)

The first field run of erikdarlingdata#1813 timed out immediately: compressed chunks
carry no btree indexes, so the V39 partial indexes cannot serve them and
the exact min() probe was a full decompress-scan cancelled by the
driver's 30s default (the probe also shipped without the explicit
timeout every sibling statement carries). The unmeasurable fail-safe
deferred correctly - but the headline behavior never engaged on exactly
the store class it was built for.

- Hypertables: floor = oldest surviving chunk's range_start from
  timescaledb_information.chunks (instant, compression-immune,
  conservative in the safe direction - prunes less, never dangles)
- Plain tables / failed conversions / empty hypertables: the exact
  V39-indexed probe, now under DeleteTimeoutSeconds
- Unmeasurable -> defer unchanged

Live test compresses the digest-carrying chunk with the product's own
settings and pins measurement, reclaim, retention, and - via a dim
placed between the chunk floor and the exact floor - WHICH path
measured; watched red by disabling the catalog branch. Full fast suite
3570 green; 12/12 live; zero warnings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant