You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
The dimension GC floor probe times out against compressed history: v39 indexes cannot serve compressed chunks #1815
Field finding from the round-5 nightly soak (first live run of #1795 / #1813), on the production field box.
Observed
The very first retention sweep after the v39 migration:
[WARN] Retention purge: could not measure query_stats's digest-carrying fact floor: Exception while reading from stream
[WARN] dimension GC deferred: a dim-feeding table's fact floor was unmeasurable this cycle; dimension content is retained until it can be measured
pg.log has the root cause: ERROR: canceling statement due to user request on SELECT min(collection_time) FROM query_stats WHERE query_text_digest IS NOT NULL OR query_plan_digest IS NOT NULL.
Two compounding causes
The probe shipped with no explicit CommandTimeout — Npgsql's 30-second default is what cancelled it. Every sibling statement in DarlingRetention carries a deliberate timeout; this one was missed.
The structural half: TimescaleDB compressed chunks do not carry regular btree indexes. The v39 partial indexes exist only on uncompressed chunks, so against months of compressed history the probe is a full decompress-scan of every compressed chunk. The Bound the dimension GC by the oldest surviving digest-carrying fact (#1795) #1813 live verification ran on an uncompressed test store and could not see this.
The unmeasurable-floor fail-safe did exactly what it was designed to do — deferred instead of guessing — so nothing is at risk in the field. But the net effect is that the headline #1795 behavior never engages on exactly the class of store it was built for (long compressed history under a coverage clamp), and the probe will time out on every sweep.
Fix shape
Hypertables: derive the floor from chunk catalog metadata — the oldest surviving chunk's range_start for each dim-feeding hypertable (one instant catalog read, compression-immune). It is conservative in the SAFE direction: every surviving fact row sits at or above its chunk's range_start, so the derived cutoff can only be older than strictly necessary, never newer — prunes less, never dangles. The ancient-orphan pathology Bound the dimension GC by the oldest surviving digest-carrying fact, not an assumed horizon #1795 targeted still dies (those sit far below any surviving chunk), and precision improves automatically as purges/backfill advance the chunk floor.
Plain PostgreSQL stores: keep the exact indexed probe (no chunks, no compression — the v39 partial index serves it as designed), now with an explicit generous CommandTimeout.
Keep the unmeasurable → defer fail-safe unchanged.
The v39 indexes stay: they serve the plain-PG path and the uncompressed tail on hypertables.
Live evidence to reproduce
A hypertable with at least one COMPRESSED chunk containing digest-carrying rows: the exact probe walks the compressed chunk; the metadata floor answers instantly. The gated live test should compress a chunk explicitly to pin this.
Fixed by #1817 (merged to dev), same day as the field finding. On hypertables the floor now comes from chunk catalog metadata - the oldest surviving chunk's range_start, one instant read, immune to compression, and conservative in the SAFE direction (every surviving fact row sits at or above its chunk's range_start, so the cutoff can only land older than strictly necessary: prunes less, never dangles). The exact V39-indexed probe stays for plain-PostgreSQL stores and unconverted tables, now under the sweep's own generous timeout instead of the driver's 30-second default. The live test now compresses the digest-carrying chunk with the product's own compression settings before measuring - the exact shape the field run died on and the one my pre-ship verification missed - and a discriminator dim placed between the chunk floor and the exact digest floor proves WHICH path measured. The affected box needs no action: the deferral was self-ending by design, and the next nightly engages the GC on its first sweep.
Field finding from the round-5 nightly soak (first live run of #1795 / #1813), on the production field box.
Observed
The very first retention sweep after the v39 migration:
pg.log has the root cause:
ERROR: canceling statement due to user requestonSELECT min(collection_time) FROM query_stats WHERE query_text_digest IS NOT NULL OR query_plan_digest IS NOT NULL.Two compounding causes
The unmeasurable-floor fail-safe did exactly what it was designed to do — deferred instead of guessing — so nothing is at risk in the field. But the net effect is that the headline #1795 behavior never engages on exactly the class of store it was built for (long compressed history under a coverage clamp), and the probe will time out on every sweep.
Fix shape
range_startfor each dim-feeding hypertable (one instant catalog read, compression-immune). It is conservative in the SAFE direction: every surviving fact row sits at or above its chunk's range_start, so the derived cutoff can only be older than strictly necessary, never newer — prunes less, never dangles. The ancient-orphan pathology Bound the dimension GC by the oldest surviving digest-carrying fact, not an assumed horizon #1795 targeted still dies (those sit far below any surviving chunk), and precision improves automatically as purges/backfill advance the chunk floor.Live evidence to reproduce
A hypertable with at least one COMPRESSED chunk containing digest-carrying rows: the exact probe walks the compressed chunk; the metadata floor answers instantly. The gated live test should compress a chunk explicitly to pin this.
🤖 Generated with Claude Code