Skip to content

Rollups serve only what they materialized: route by coverage, and back fill behind a disk preflight - #1788

Merged
erikdarlingdata merged 17 commits into
devfrom
feature/1759-coverage-routing-and-backfill
Jul 28, 2026
Merged

erikdarlingdata merged 17 commits into
devfrom
feature/1759-coverage-routing-and-backfill

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Jul 28, 2026 •

Copy link
Copy Markdown
Owner

Closes #1759 (Phases 1 and 2). Phase 0 — the held-paused observability WARN lines — already shipped in #1762.

One PR, two commits, and why

Phase 2 hard-depends on Phase 1's coverage probe: the backfill's resume point, its convergence check and the router's fallback all read the same measured floor, and splitting them would have meant a stacked PR against a protected branch on the night before a release cut. The two commits are cleanly separated and reviewable in order — Phase 1 is read-side only and could ship alone; Phase 2 is an operator verb nothing invokes by accident.

Phase 1 — coverage-aware routing (read-side, zero disk risk)

RetentionTierRouter routed by AGE alone, and nothing fell back to raw. On a store whose rollups were created WITH NO DATA over pre-existing history, every window older than the rollup's materialized floor read empty while raw still held every row — and raw still held them precisely because the #1680 arming gate had held its purge closed for the same reason.

RollupCoverage (a companion to #1665's RollupAvailability) probes each rollup's min(bucket) plus each rolled raw table's min(collection_time) in one round trip, and the router degrades a window to whichever tier is measured to reach furthest back.

The hard half is not falling back when raw would return LESS. On a healthy store with armed purges, raw keeps ~4 days against the rollups' weeks, so a window older than every floor is the normal case there — a naive "window predates the floor → use raw" rule would send every long window to a 4-day table. The rule is therefore comparative: a tier is abandoned only on a positive measurement that a lower tier is deeper, which is exactly the held-purge signature #1759 describes and is silent everywhere else. Half the routing truth table exists to hold that line, and there is a live-PG test for it specifically.

Unknown coverage is inert by construction — a failed probe, a plain-PostgreSQL store and a partially-built one all produce nulls, and nulls never move a window, so the pre-#1759 behaviour is reproduced exactly.

Other things this needed:

  • All six production routing call sites gated, enumerated: ComposeSourceRouter.cs:178 (one site, routing dynamically across all three catalog tables), Mcp/DarlingHealthReader.cs:210, ViewerDataService.DailySummary.cs:56, and three in ViewerDataService.FinOps.Workload.cs (:205 on the db-grain pair, :363 and :429 on the query-grain pair). (An earlier draft of this body said seven — a miscount, corrected. Six is the verified number; the guard's rot detector now pins that floor.) The MCP get_daily_health reader is not incidental: it answers the same question as the viewer's Performance Calendar off the same shared SQL, so gating one and not the other would have had them report different query counts for the same day on exactly the affected stores, with no way to tell which was right.

    The source-parsing guard that keeps a future reader from being added un-gated is deliberately two-sided — a positive match on a real coverage lookup (.For() and an explicit rejection of TierCoverage.Unknown. Its first cut was one-sided ("does the argument list mention coverage?") and was holed: "overage" is a substring of the type name TierCoverage, so a reader passing TierCoverage.Unknown satisfied it. That form compiles and reads as deliberate, which makes it the most probable way Rollup CAGGs serve only materialized buckets: old windows read empty and raw purges stay held #1759 returns — not someone dropping an argument, which a compile-shaped review also catches, but someone reaching for the inert value because it satisfied the signature. Both forms are now pinned as mutations (table below). Tightening it also exposed that the scan was matching <see cref="…Resolve(…)"/> inside doc comments and running that phantom match forward into real code; the one-sided predicate had hidden it, because a cref contains the type name. Comments are stripped before scanning now.

  • Coverage expires even where availability does not. Availability is cached permanently once complete, because a created aggregate is never dropped — true of existence, false of coverage, which moves backwards on a backfill and forwards on a retention drop. Keeping the AllPresent shortcut would have pinned a pre-backfill floor for the life of the process, so an operator who had just backfilled would keep getting raw fallbacks until a restart. The existing 5-minute reprobe interval now applies unconditionally, in both the composer's per-datasource cache and the viewer's.

  • The partial-window notice now prefers the MEASURED floor to the retention span. On these stores the purges are held, so raw holds months rather than its nominal ~4 days; left assuming, the notice would have stamped "older points are not included" on precisely the coverage-fallback panels that are in fact complete — a false alarm introduced by the fix. It falls back to the span when nothing is measured, which reproduces the old text.

The false premise is retired. TimescaleSupport's claim that "real-time aggregation is opted into… correct to query for any window immediately" was wrong twice over — 2.13+ defaults materialized_only to TRUE, and even ON the watermark is a hard partition, so un-materialized history is served by neither branch. The materialized_only test pin stays, now load-bearing and for the true reason: RollupCoverageProbeSql and RetentionArmSafetySql both read min(bucket) to mean "the oldest bucket MATERIALIZED", and unioning the raw branch in would make an empty materialization report raw's own oldest row as the rollup's floor — the router would believe coverage it does not have, and the arming gate would arm a purge over history nothing else holds. The cold-start caveat prescribing a manual backfill no store ever received is retired too, replaced by a pointer to the verb.

Phase 2 — --backfill-rollups, an operator verb with a disk preflight

The rejected option ("just open the gate") is not resurrected. The arming gate is all-or-nothing, so a store with a year of raw must materialize the whole history before the first purge arms and reclaims anything: peak disk comes before any relief. At service start, on the worst-affected stores, that is a plausible disk-exhaustion event.

  • Preflight that refuses with numbers. The estimate is calibrated from what the rollup has already materialized (bytes per bucket × buckets to add) — every affected store has that sample, since its refresh policy has been materializing a trailing 3-day window all along. With no sample it bounds from raw's own size and says so, deliberately erring high: refusing a backfill that would have fit is recoverable, filling a production volume is not. Requires the estimate + 25% headroom + a 10 GB reserve. Free space is measured on the volume the store reports (current_setting('data_directory')), so a store on another host refuses rather than measuring this machine's disk. The refusal names the shortfall and both real options — including that waiting is safe, since nothing is being lost while the purges are held.
  • Sliced one source chunk at a time, NEWEST first (see the BLOCKING-2 fix above — the direction is a data-safety property). The issue's own API research is right that a manual refresh already batches internally, so this is supervision rather than re-implementation: progress on a multi-hour run, a resume point, and a lock window short enough not to sit across a compression job (the Compression waits for the next fixed tick, so a just-closed chunk can sit uncompressed most of a day (81 GB observed); deadlock correlation to watch #1778 watch).
  • Convergence read from DATA. A batch-cap stop is logged server-side and completely silent to the client, so the calls returning is no evidence. Only a measured shortfall escalates to the forced form — the one thing that repairs a hole left by an interrupted pass, whose invalidation records a plain refresh now skips straight over.
  • Idempotent and resumable — every pass re-plans from the measured floor; a completed one converges to a no-op.
  • Arms nothing. Coverage is the whole job. The arming gate already self-heals at startup, and duplicating that decision in a second place is how the one thing protecting this data stops being the one thing.
  • --dry-run prints the plan, the estimate and a duration budget at the ~16 MB/s measured on this host class, and stops.

Concurrency / locking. refresh_continuous_aggregate takes no lock that blocks writers on the source hypertable, so collection keeps running; it cannot run inside a transaction, so each slice is its own statement and partial progress survives an abort. What it does contend with is the compression policy on the same chunks, which is why slices are one chunk wide.

The operator contract, verbatim

Every line the verb tells an operator is reproduced here rather than paraphrased, because these are the contract and a reviewer should be able to read them without opening the source. Both are pinned by BackfillRollupsVerb_DryRunsThenBackfills_AndReportsCoverageItActuallyReached, so editing them breaks a test rather than silently voiding the contract.

On completion — the next step, the confirmation to look for, and the one interaction the operator has to know about: the hourly rollups carry their own 21-day retention policy, already armed on these stores, which will trim the coverage a run just built when it next fires (measured cadence: schedule_interval = 1 day, see #1790):

NEXT: restart the PerformanceMonitor Darling service. The arming gate checks coverage at
startup and releases the held retention policies by itself — there is no arming step here and
nothing to run by hand. The startup log line reading
  'N/N retention policies in place, N armed, 0 held paused pending backfill'
is the confirmation; the first purge then reclaims the raw tables in one pass.

Do not delay the restart. The hourly rollups carry their OWN 21-day retention policy, already
armed on these stores, which will trim the coverage this run just built when it next fires
(roughly daily). Restarting now is what lets the raw policies arm off that coverage first. If
the trim wins the race nothing is lost — raw is still held — and re-running this verb rebuilds it.

And the sibling line on the nothing-to-do path (DarlingCliCommands.cs:2187), which carries the same self-heal promise for the store that is already covered — including a re-run after a successful pass, which is the most likely way an operator sees this verb a second time:

Every rollup already covers its raw table. Any retention policy still held will arm itself on the next service start.

Two BLOCKING review defects, both proven by execution, both fixed

1. The preflight under-estimated by rows-per-bucket. The probe counted ROWS (count(*)) and the arithmetic divided that into bytes to get a per-BUCKET figure, then multiplied by a BUCKET count. A rollup holds one row per (server, database, query_hash, sql_handle, bucket), so the units differ by the distinct queries seen per hour — 205x under on a 200-row/hour store, worse on a real fleet box. That is the one direction this preflight exists to prevent, and the code's own comment said so. Fixed with count(DISTINCT bucket).

The suite could not have caught it, which mattered as much as the bug: both live fixtures seeded one row per bucket, making rows and buckets the same number and the error factor exactly 1. That is not a small fixture, it is a fixture of a shape the product never meets. They now seed 12 distinct queries per bucket, and a live test asserts the probe returns a bucket count by ratio rather than by string — so the semantics are pinned, not the spelling.

2. An interrupted backfill left a hole, reported DONE, and the arming gate then armed a purge over it. min(bucket) was the resume point, the completion verdict, and what the #1680 gate reads before letting the 4-day raw purge drop chunks. Slices ran oldest-first, so the first slice drove that single number to its final value and every later slice was invisible to all three. Proven by execution: killed at slice 32 of 43, 47 of 264 buckets missing, re-run printed "nothing to do — DONE" and exited 0, and the gate then armed over a window whose only surviving copy was the raw rows about to be dropped. The slice-failure path reached the same state, and the "Safe to interrupt" line was false.

Fixed by reversing the slice order, not by adding per-slice bookkeeping. Descending makes floor <= raw_oldest imply completeness by construction: the floor can only reach the bottom once the last, oldest slice has run, in any interleaving. An interrupted run leaves it visibly short — a truthful SHORT instead of a false DONE — and the gate stays closed because it reads the same honest number. Mid-slice kills need no extra state either: a refresh commits per batch newest-first, so the floor lands inside the killed slice and re-planning from it re-covers the remainder. I chose this over per-slice verification because it removes the failure mode rather than detecting it — there is no interleaving left in which a hole can be reported as coverage, so nothing depends on a probe being remembered.

Plus the non-blocking sanity clamp. One stray epoch-era row produced a 739,825-slice plan that burned CPU indefinitely. Plans past a 10-year ceiling now REFUSE and name the offending timestamp — that is corruption someone has to go and look at, not history to back fill. A refusal is deliberately not a skip: [REFUSED] on stderr, DONE blocked, non-zero exit, because burying it among the [OK] lines is exactly how it would be missed.

Three findings from the gated live leg

None were visible by reading; all three are product fixes, not test workarounds.

  1. 55P03 concurrent refresh. The verb runs while the service is UP and the aggregate's own refresh policy lands on top of a slice. Reproduced immediately: EnsureContinuousAggregatesAsync attaches the policy, the policy runs at once, the next slice failed. Retried, bounded, transient-only — every other SQLSTATE still fails fast.
  2. 22023 "refresh window too small". A day-wide slice's ragged tail is narrower than one bucket for a daily rollup, and the range end is a coverage floor or "now", so it lands mid-bucket most of the time. It aborted the whole daily tier. The planned range now closes on a bucket boundary as well as opening on one.
  3. The per-slice convergence check was measuring the wrong thing. "Did the floor reach this slice's start?" fired on the first slice of every run, because a slice's range can legitimately hold no source rows — raw's oldest row lands partway into the first slice, and a collection gap does the same mid-run. A global min(bucket) cannot answer a question about one slice. Convergence is judged once, at the end, against raw's oldest row.

Testing

Full solution rebuild: 0 warnings, 0 errors. Darling suite: 3723 passed, 0 failed (7 skipped — gated on a live SQL Server and the PG-runtime fixture, neither present here).

Gated live legs ran against a real PostgreSQL 18.4 + TimescaleDB 2.28.1 store stood up from the repo's pg-runtime.zip, on a store built into the exact broken shape: pre-existing history, rollup created WITH NO DATA, floor well after raw's.

  • The defect measured, not assumed: a window below the floor returns 0 rows from the rollup while raw holds rows for the same window.
  • Phase 1 routes that window to raw; after the backfill it routes back to the rollup.
  • The healthy shape (raw shallower than the rollup) does not fall back — with the row counts showing why.
  • The verb driven end to end through a bring-your-own darling.json, not a re-implementation of its loop: dry run changes nothing, the real run's DONE claim is re-verified against the store per rollup, and a second run reports nothing to do.
  • The interrupted-backfill acceptance — the reviewer's data-loss proof inverted. Kill at two thirds of the slices, then assert the partial run does not look complete, the re-plan still has work, the real arming gate refuses (read from the job catalog, not by re-evaluating its predicate), then resume and assert zero missing buckets measured against raw itself rather than against our own arithmetic — and only then does the gate arm.
  • The verb's refusal branch end to end, which previously had only its pure arithmetic covered: a single poisoned timestamp, [REFUSED] on stderr naming it, no DONE, exit 1.
  • The arming gate then arms query_stats' held retention policy by itself, through EnsureRetentionPoliciesAsync — the real seam, not its predicate.

Mutation table, each verified RED:

Mutation Caught by
Floor check dropped 7 routing tests; the held-store divergence pin goes Raw → Daily
"Raw is deeper" guard removed (naive floor check) 8 tests, incl. the healthy-store no-regression legs
(A) A routing reader passing TierCoverage.Unknown — compiles, reads as deliberate, and is the likely way back Guard: ViewerDataService.DailySummary.cs:57 (passes TierCoverage.Unknown — routes on NO evidence)
(B) A routing reader with the coverage argument removed entirely Guard: ViewerDataService.DailySummary.cs:57 (passes no coverage lookup)
(C) Row-count estimator restored (count(*)), against the R>1 fixture Unit pin and the live ratio guard — the old 1-row-per-bucket fixture caught neither
(D) Ascending slice order restored The live acceptance reproduces the defect verbatim: "an interrupted backfill reports coverage … at or below raw's oldest — it looks COMPLETE", plus two unit pins
Preflight bypassed 3 preflight theory rows
Convergence made call-based over a silently under-covering plan Verb reports INCOMPLETE + names the exact shortfall; live test red
Hierarchical order inverted The order pin and the live verb (dailies genuinely short)

Known limits of the coverage model, recorded not assumed away

The floor is a depth measure, not a completeness guarantee, and RetentionTierRouter's docs now say so explicitly. Two shapes slip through a one-dimensional floor: a mid-window materialization hole is invisible to it (a service down longer than the refresh policy's 3-day start_offset resumes at now-3d and never backfills the skipped interval, while min(bucket) goes on reporting the original deep floor), and a window straddling the floor is served partially with no signal to the caller (returning the tier is still the right choice — raw would return less on a healthy store — but the part below the floor is missing).

Both are accepted here because both are strictly better than the age-only routing they replace, which served the entire window as EMPTY in exactly these cases. The docs paragraph exists to block one specific inference: "the coverage gate passed, therefore the result is complete." That belief-shape is how #1759 survived as long as it did — the previous comment asserted the view was "correct to query for any window immediately", it read as reasonable, it was pinned by a test, and nobody re-derived it for months.

Tracked as #1791 with both candidate closures scoped (two-dimensional coverage with gap detection; a partial-coverage notice through the existing CombineNotices machinery). Neither is cut-night material — the first has a real cost question on a 5-minute probe cadence, the second is a compose-contract change.

Follow-ups filed

Interaction with the other lanes in flight

Verified rather than assumed: #1772 (delta-seed time bound) and #1768/#1767 (payload dimensions) touch neither the routing ladder nor the rollups' materialization — the CAGGs never carried the payload columns, so the dimension work is invisible here, and the delta seeder writes raw rows the coverage probe only ever reads a min() from. Branched off origin/dev at 50012266, which already contains both. Full-suite green above is over the merged tree.

🤖 Generated with Claude Code

erikdarlingdata and others added 3 commits July 27, 2026 23:59
A continuous aggregate created WITH NO DATA over pre-existing history serves
only what it materialized. Real-time aggregation cannot rescue the rest: the
watermark is a hard partition (materialized below UNION ALL raw at-or-above),
so history below it that was never materialized is served by NEITHER branch.
Every rollup's refresh policy starts 3 days back, so on a store that existed
before its rollups the materialized span begins at creation-minus-3-days and
never reaches further back on its own.

Age-only routing therefore sent old windows to a rollup that answered with
silence while raw still held every row -- and raw still held them precisely
because the #1680 arming gate had held its purge closed for the same reason.

RollupCoverage probes each rollup's min materialized bucket plus each rolled
raw table's oldest row, and RetentionTierRouter degrades a window to the tier
measured to reach furthest back.

The hard half is NOT falling back when raw would return LESS. On a healthy
store raw keeps ~4 days against the rollups' weeks, so a window older than
every floor is the NORMAL case there and a naive floor check would send every
long window to a 4-day table. The rule is comparative: a tier is abandoned
only on a positive measurement that a lower tier is deeper. Unknown coverage
(failed probe, plain PostgreSQL, partial build) is inert by construction.

Also:
- All 7 production routing readers gated, including the MCP get_daily_health
  reader, which answers the same question as the viewer's calendar off the
  same SQL -- gating one and not the other would have them disagree about how
  many queries ran on a given day, on exactly the affected stores. A
  source-parsing guard keeps a future reader from being added un-gated.
- Coverage expires on the existing 5-minute reprobe interval even when
  availability does not. Availability is permanent once complete (a created
  aggregate is never dropped); coverage MOVES -- backwards on a backfill,
  forwards on a retention drop -- so the AllPresent permanent-cache shortcut
  would have pinned a pre-backfill floor for the life of the process.
- The partial-window notice now prefers the MEASURED floor over the retention
  span. On these stores the purges are held, so raw holds months rather than
  its nominal ~4 days, and the assumed span would have stamped "older points
  are not included" on precisely the fallback panels that are complete.
- Retires the false premise at the CAGG definitions ("real-time aggregation
  opted into... correct to query for any window immediately") and the
  cold-start caveat prescribing a manual backfill no store ever received. The
  materialized_only test pin stays, now load-bearing and for the true reason:
  the coverage probe and the arming gate both read min(bucket) to mean "the
  oldest bucket MATERIALIZED", which the real-time union branch would break.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
With reads already correct (Phase 1), materializing the rollups becomes a
capacity operation rather than a correctness emergency -- which is what lets
it be preflighted and operator-triggered instead of run at startup.

It is NOT a startup step, and that is the whole shape of it. The #1680 arming
gate is all-or-nothing, so a store with a year of raw has to materialize the
WHOLE history before the first purge arms and reclaims anything: peak disk
comes BEFORE any relief. Doing that automatically at service start, on the
exact stores worst affected (one already down to ~150 GB free), is a plausible
disk-exhaustion event.

- DISK PREFLIGHT that refuses with numbers. The estimate is CALIBRATED from
  what the rollup has already materialized (bytes per bucket x buckets to
  add), which every affected store has, since its refresh policy has been
  materializing a trailing 3-day window all along; with no sample it bounds
  from raw's own size and flags itself as a bound. Requires the estimate plus
  25% headroom plus a 10 GB reserve. Free space is measured on the volume the
  STORE says it lives on (current_setting('data_directory')), so a store on
  another host refuses rather than measuring this machine's disk. The refusal
  names the shortfall and both real options, including that waiting is SAFE --
  nothing is being lost while the purges are held.
- SLICED one source chunk at a time, oldest first. The engine already batches
  within a call, so this is supervision, not re-implementation: progress on a
  multi-hour run, a resume point, a lock window short enough not to sit across
  a compression job (#1778), and no unbounded transaction.
- CONVERGENCE READ FROM DATA. A refresh that stops on its internal batch cap
  logs server-side and returns success to the client, so the calls returning
  is no evidence. Only a measured shortfall escalates to the forced form,
  which is the one thing that repairs a hole left by an interrupted pass.
- IDEMPOTENT and resumable: every pass re-plans from the measured floor.
- ARMS NOTHING. Coverage is the whole job; the arming gate already self-heals
  on the next service start, and duplicating that decision in a second place
  is how the one thing protecting this data stops being the one thing.

Three findings from the gated live leg (PostgreSQL 18.4 / TimescaleDB 2.28.1),
none of which were visible by reading:

1. 55P03 concurrent refresh. The verb runs while the service is UP, and the
   aggregate's own refresh policy lands on top of a slice. Retried, bounded,
   transient-only -- every other SQLSTATE still fails fast.
2. 22023 "refresh window too small". A day-wide slice's ragged tail is
   narrower than one bucket for a DAILY rollup, and the range end is a
   coverage floor or "now", so it lands mid-bucket most of the time. It
   aborted the whole daily tier. The planned range now closes on a bucket
   boundary as well as opening on one.
3. The per-slice "did the floor reach this slice's start?" check was measuring
   the wrong thing -- a slice's range can legitimately hold no source rows
   (raw's oldest row lands partway into the first slice; a collection gap does
   the same mid-run), so it fired on the first slice of every run. A global
   min(bucket) cannot answer a question about one slice. Convergence is judged
   once, at the end, against raw's oldest row.

Tested with the verb driven end to end against a real store through a
bring-your-own darling.json, not a re-implementation of its loop: dry run
changes nothing, the real run's DONE claim is re-verified against the store,
and a second run reports nothing to do. Every guard verified RED by mutation:
preflight bypassed, convergence made call-based over a silently under-covering
plan, and the hierarchical order inverted (caught by the pin AND by the live
verb, which reported the dailies genuinely short).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
erikdarlingdata and others added 2 commits July 28, 2026 00:16
…ntee

Comment-only. Names the two shapes a one-dimensional floor cannot see, and
blocks the inference that would let them be forgotten again.

A mid-window HOLE is invisible: coverage is one number, so it cannot see a
gap above itself. A service down longer than the refresh policy's 3-day
start_offset resumes at now-3d and never backfills the skipped interval,
while min(bucket) goes on reporting the original deep floor -- so the window
routes to the tier and is served as complete with the gap inside it.

A window STRADDLING the floor is served partially with no signal. Returning
the tier is the correct CHOICE there (raw would return less on a healthy
store), but the part below the floor is missing and nothing says so.

Both are accepted: both are strictly better than the age-only routing they
replace, which served the ENTIRE window as empty in exactly these cases. What
the paragraph exists to block is the belief "the coverage gate passed,
therefore the result is complete" -- which is precisely how #1759 survived as
long as it did. The previous comment asserted the view was "correct to query
for any window immediately", it read as reasonable, it was pinned by a test,
and nobody re-derived it for months. A one-dimensional measure invites the
same inference in a new place.

Both gaps and their candidate closures are tracked in #1791 rather than
assumed away here.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The guard standing between #1759 and its own return could be satisfied by the
most probable way back.

It checked that a Resolve call's argument list contained the substring
"overage" -- which is a substring of the TYPE NAME TierCoverage. So a reader
passing TierCoverage.Unknown, the canonical "route on no evidence" value,
SATISFIED the guard. That form compiles and reads as deliberate, which is
exactly what makes it likely: not someone dropping an argument (the form this
had been verified against, and one a compile-shaped review also catches), but
someone reaching for the inert value because it satisfied the signature.

The predicate is now two-sided -- a positive match on a real coverage lookup
(.For() AND an explicit rejection of TierCoverage.Unknown. Either half alone
is holed: "mentions coverage" is satisfied by the type name, and "does not say
Unknown" is satisfied by passing nothing at all. The failure message names
which of the two it found, since they need different fixes.

Tightening it surfaced a second flaw the one-sided predicate had been hiding:
the scan matched <see cref="...Resolve(...)"/> inside DOC COMMENTS and, since
the pattern spans newlines looking for the closing ");", ran that phantom
match forward into whatever real code followed. It passed silently because a
cref contains the type name -- and it inflated the rot detector's count.
Comments are stripped before scanning now, preserving line numbers so reports
still point somewhere openable.

Rot detector raised 5 -> 6, the verified production call-site count: the
composer (one site, routing dynamically across all three catalog tables), the
MCP daily-health reader, the viewer's calendar, and three FinOps readers. A
legitimate removal should lower it deliberately rather than be absorbed.

Both forms verified RED against a green baseline, each naming its own cause.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
erikdarlingdata and others added 3 commits July 28, 2026 00:45
Comment-only. #1776 landed a hygiene guard on dev while this branch was open:
every class reaching the shared DARLING_TEST_PG store must either serialize
with [Collection("live-postgres")] or say in writing why it does not.

This class mints its own scratch database, so it cannot race the shared store
and serializing it would be pure slowdown -- it is the exemption case, and it
now carries the "#1776 own-store" marker the guard looks for. It has to own
its store regardless: it creates continuous aggregates and retention policies
the shared fixture must never inherit from a test.

Found by test-merging origin/dev rather than trusting the merge status. The
merge is TEXTUALLY clean on everything but a keep-both CHANGELOG link-ref
block -- git reported no conflict here at all, because a new guard on one side
and a new class on the other do not overlap as text. Only building and running
the merged tree surfaced it. Fixing it here rather than at merge time keeps
the eventual resolve to the CHANGELOG alone.

Verified on the merged tree: LivePostgresCollectionHygieneTests passes and the
full suite is green (3546 passed, 0 failed). This branch alone: 3537 passed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Resolved to get CI to run at all, not as an arm-time step: a CONFLICTING PR
has no merge ref, and pull_request workflows build the merge ref -- so Build
and Claude Auto Review could not start on this PR while the conflict stood,
while other branches' runs kept firing normally. That is what made this look
like an Actions outage for ~25 minutes.

CHANGELOG.md was the only conflict, and only in the link-ref block. Resolved
KEEP-BOTH with no re-sorting: this branch's #1665/#1788 refs and dev's
#1781/#1783/#1786/#1792 refs all retained, in the order they appeared. No
entry text on either side was touched.

Everything else auto-merged. Verified rather than assumed, because the
dangerous case here is textual cleanliness hiding a semantic break: full
solution rebuild 0 warnings / 0 errors, Darling suite 3549 passed / 0 failed.
The one real collision this class of merge produced -- #1776's new live-store
hygiene guard versus this branch's new own-store live test -- was found the
same way and fixed in 59f877a before this merge.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ill hole

BLOCKING 1 -- the preflight under-estimated by rows-per-bucket.
The probe counted ROWS (count(*)) and the arithmetic divided that into bytes
to get a per-BUCKET figure, then multiplied by a BUCKET count. A rollup holds
one row per (server, database, query_hash, sql_handle, bucket), so the two
units differ by the distinct queries seen per hour -- measured 205x under on a
200-row/hour store, worse on a real fleet box. Under-estimating is the single
direction this preflight exists to prevent, and my own comment said so.
count(DISTINCT bucket) fixes it.

The suite could not have caught it: BOTH live fixtures seeded one row per
bucket, making rows and buckets the same number and the error factor exactly
1. That is not a small fixture, it is a fixture of a shape the product never
meets. They now seed 12 distinct queries per bucket, and a live test asserts
the probe's count is a bucket count by ratio rather than by string.

BLOCKING 2 -- an interrupted backfill left a hole, reported DONE, and the
arming gate then armed a purge over it.
min(bucket) is the resume point, the completion verdict, AND what the #1680
gate reads before letting the 4-day raw purge drop chunks. Slices ran
OLDEST-FIRST, so the FIRST slice drove that one number to its final value and
every later slice was invisible to all three. Proven by execution: killed at
slice 32/43 -> 47/264 buckets missing -> re-run printed "nothing to do --
DONE" and exited 0 -> the gate armed over a window whose only surviving copy
was the raw rows about to be dropped. The slice-failure path reached the same
state. My "Safe to interrupt" line was false.

Fixed by REVERSING the slice order rather than adding per-slice bookkeeping.
Descending makes floor <= raw_oldest imply completeness BY CONSTRUCTION, in
any interleaving: the floor can only reach the bottom once the last, oldest
slice has run. An interrupted run leaves it visibly short, which is a truthful
SHORT rather than a false DONE, and the gate stays closed because it reads the
same honest number. Mid-slice kills need no extra state: refresh commits per
batch newest-first, so the floor lands inside the killed slice and re-planning
from it re-covers the remainder.

Chose this over per-slice verification because it removes the failure mode
instead of detecting it -- there is no interleaving left in which a hole can
be reported as coverage, so nothing depends on a probe being remembered.

Also folded in the sanity clamp: one stray epoch-era row produced a
739,825-slice plan that burned CPU indefinitely. Plans past a 10-year ceiling
now REFUSE and name the offending timestamp -- corruption someone must look
at, not history to back fill. A refusal is deliberately not a skip: it prints
[REFUSED] to stderr, blocks DONE, and exits non-zero, because burying it in
the [OK] lines is how it would be missed.

Every fix verified RED by mutation against a green baseline:
- count(*) restored -> unit pin AND the live ratio guard both fail
- ascending order restored -> the live acceptance reproduces the exact defect
  ("an interrupted backfill reports coverage ... at or below raw's oldest --
  it looks COMPLETE"), plus two unit pins
New coverage: the reviewer's proof inverted (kill at 2/3, assert the partial
run does not look complete, the re-plan still has work, the REAL arming gate
refuses, then resume and assert ZERO missing buckets measured against raw
itself, then the gate arms), and the verb's refusal branch end to end, which
previously had only its pure arithmetic covered.

Full solution rebuild 0 warnings / 0 errors; Darling suite 3723 passed / 0
failed against live PostgreSQL 18.4 + TimescaleDB 2.28.1; backfill tests run
three times for flake.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
erikdarlingdata and others added 9 commits July 28, 2026 01:28
Resolved promptly rather than at arm time: a DIRTY PR gets no pull_request CI
at all (the merge ref is uncomputable), so leaving it dirty would strand the
two BLOCKING fixes in be0d2c6 without check evidence for either review seat.

CHANGELOG.md auto-merged this time -- no conflict, nothing to hand-resolve.
Adopted proof run and clean: `git diff origin/dev -- CHANGELOG.md` shows only
the two link-refs this branch ADDS (#1665, #1788), with no line removed and
none moved.

Brought in from dev: DarlingRetention.cs and PayloadDimensionLiveTests.cs,
neither of which this branch touches.

Verified rather than trusted, because MERGEABLE is not "the merge is safe" --
a guard landing on one side and a new class on the other do not overlap as
text but do collide semantically, which is exactly how #1776's hygiene guard
caught this branch's own live test earlier tonight. Full solution rebuild
0 warnings / 0 errors; Darling suite 3724 passed / 0 failed against live
PostgreSQL 18.4 + TimescaleDB 2.28.1, so the six #1759 live legs ran rather
than skipping.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Both are properties the fix DEPENDS on that nothing was enforcing, raised by
verb-reviewer's re-verification.

1. refresh_newest_first. The no-extra-state mid-slice resume argument rests on
   a killed slice leaving its NEWEST batches committed, so the floor lands
   inside that slice and the next run's top slice re-covers the remainder.
   That is true only because refresh_newest_first is TRUE -- which is the 2.28
   DEFAULT taken when options is NULL. Verified the verb passes no 5th
   argument on either form, and it now says so at the call site rather than
   leaving the dependency invisible. The pin fails the moment anything starts
   passing options, at which point the assumption has to be made explicit
   (pass refresh_newest_first by name) instead of inherited from a value
   somebody else chose. Asserted on the forced form too: it already takes a
   4th argument and is the likeliest place a 5th gets appended.

2. Bucket alignment across the reversal (P2b). The alignment property was
   established against ASCENDING slicing, where the ragged remainder sat at
   the TOP of the range. Reversing moved it to the BOTTOM -- a different
   slice, produced by a different branch of the loop (the clamp to `from`
   rather than the clamp to `to`) -- so the property had to be re-established
   rather than assumed to have carried. Now checks BOTH ENDS of every slice
   against the bucket grid, not just the width: a slice one bucket wide but
   half a bucket out of phase is still refused with 22023. Plus a pin that the
   remainder really is at the bottom, which would catch a silent revert to
   ascending putting it back on top and re-opening the false-DONE path.

Both verified RED against a green baseline, on mutations that COMPILE:
appending `false, $3::jsonb` to the refresh call, and restoring ascending
order. The first attempt at the options mutation did not compile and the test
ran against a stale binary reporting green -- worth naming, because a mutation
that fails to build is indistinguishable from a guard that works.

Full solution rebuild 0 warnings / 0 errors; Darling suite 3725 passed / 0
failed against live PostgreSQL 18.4 + TimescaleDB 2.28.1.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The reviewer's premise check was right and the comment was wrong. It said
refresh_newest_first "is on" as though it were configured -- but there is no
such GUC on 2.28.1 (not in pg_settings). It rides the options jsonb parameter,
which this verb never passed, so the mid-slice resume guarantee rested on an
UNSTATED SERVER DEFAULT.

It holds today, and a real mid-slice kill measured zero holes above the floor.
That is exactly what makes it worth fixing rather than leaving: this PR's
original defect came from the same class of trust. materialized_only was also
assumed, also read as reasonable, also pinned by a test, and was wrong for
months. An unstated default that happens to be right is not a contract.

So the call now passes options => {"refresh_newest_first": true} explicitly,
force becomes positional to reach it, and the comment states what the code
ASSERTS rather than what it hopes. Only newest-first is set: batch size and
the batch cap stay at the defaults the issue's API research established are
already the right shape.

Degrades honestly rather than hard-failing. options arrived in 2.21 and
nothing gates a bring-your-own store's version, so a pre-2.21 store raises
42883 on the 5-argument call. RefreshAsync falls back to the 3-argument form
ONCE and says so out loud -- on such a store the guarantee genuinely IS an
inherited default, and telling an operator a contract is held when it is not
would be the same failure in a new place. Branched on SQLSTATE, never the
message.

The pin is inverted accordingly: it previously asserted no options were
passed, and its own doc said "anything that starts passing options must pass
refresh_newest_first explicitly and turn this into a pinned property rather
than an assumption". That is what happened, so it now asserts the contract IS
carried, on both the plain and forced forms, plus the degraded 3-argument
shape and the SQLSTATE it keys on.

Verified against the real store rather than assumed: the full interrupted-
backfill acceptance re-ran on live PostgreSQL 18.4 / TimescaleDB 2.28.1 with
the explicit option in place -- 2.28.1 accepts it and the kill/resume/zero-
missing-buckets/gate-arms chain still holds. Mutation (withOptions default
flipped back to false) verified RED on a COMPILING build.

Full solution rebuild 0 warnings / 0 errors; Darling suite 3725 passed / 0
failed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Test-only hardening, prompted by a CI flake I could not honestly call
unrelated.

The darling-pg leg failed once on 28987fd with 15 tests dying inside one
second on "existing connection was forcibly closed" -- a server-level event,
not a logic failure, and mostly in shared-store classes this branch does not
touch. A re-run of the same SHA passed, so it is not deterministic and the
explicit-options change did not cause it. But a green retry proves only that
it is not deterministic; it does not make my tests innocent, and one of the
casualties WAS mine.

What I did introduce is a genuinely destabilising pattern: these tests ARM
real retention policies, because whether the #1680 gate DECIDES to arm is the
thing under test. An armed policy means TimescaleDB's scheduler will launch a
background worker to run drop_chunks -- against a scratch database
ScratchPostgres is about to DROP ... WITH (FORCE). Leaving a worker running
against a vanishing database is a hazard whether or not it caused this
particular flake.

So both arming tests now unschedule their jobs once the assertions have read
the catalog. That costs the tests nothing: what is under test is the gate's
DECISION, already measured by then, and nothing needs the purge to actually
execute. Best-effort, because a cleanup failure that fails a passing test
inverts the signal.

Deliberately NOT a retry wrapper. The recorded lesson from the earlier
live-PG flake is to harden the class's lifecycle rather than paper over it,
and a retry here would have hidden exactly the pattern worth removing.

Full solution rebuild 0 warnings / 0 errors; Darling suite 3725 passed / 0
failed against live PostgreSQL 18.4 + TimescaleDB 2.28.1.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Resolve-only push during an active review window. CHANGELOG auto-merged with
no conflict; adopted proof clean (only this branch's #1665/#1788 link-refs
added, none removed or moved).

Dev brought in the dims pair plus TimescaleSupport.cs, which this branch also
edits -- so verified rather than trusted, per MERGEABLE-is-not-merge-safe:
full solution rebuild 0 warnings / 0 errors, Darling suite 3736 passed / 0
failed against live PostgreSQL 18.4 + TimescaleDB 2.28.1 with the live legs
actually running.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two decisions, both reversals of my own previous commit.

REVERTED the explicit options argument. The compatibility matrix corrected
both the prescription and my implementation of it: options arrived in
TimescaleDB 2.21, nothing in this product gates a bring-your-own store's
version, and my degrade-on-42883 fallback re-attempted the 5-argument form on
EVERY slice -- so a 2.18-2.20 store would have eaten a failed round trip per
slice, for a backfill that already works there. That is a real cost paid to
guard an engine flip that has not happened. The no-options shape is restored;
the comment stays corrected, because the original lie was calling an inherited
ENGINE DEFAULT a configured setting, and that is still worth saying plainly.

REPLACED it with a guard on the BEHAVIOUR. Every other test in this file
kills BETWEEN slices, so the mid-slice premise -- a kill inside a slice leaves
the floor inside the cancelled range with no holes above it -- had nothing
watching it; its only evidence was hand-executed kills in a transcript. The
new pin cancels a wide refresh mid-flight and asserts exactly that. It never
mentions refresh_newest_first, so it holds on any engine including BYO stores
that would reject the option outright, and a bundled-runtime bump re-runs it
automatically.

THE FIXTURE IS THE WHOLE TEST, and my first two attempts were vacuous. The
ensure sweep also attaches a refresh policy, TimescaleDB runs a new policy's
first check IMMEDIATELY, and that policy materializes a trailing window
CONTIGUOUSLY -- which satisfies both assertions on its own, whatever the
cancelled refresh did. An oldest-first mutation PASSED twice: once with the
policy live, and again after unscheduling it post-creation, which is too late.
The aggregate is now created from its own DDL with NO policy attached, so the
only materialization in the database is the one being cancelled. Only then did
the mutation go red -- 2832 missing buckets, correctly diagnosed as "this
engine did NOT commit its batches newest-first".

Also reordered the assertions so a flipped engine is diagnosed as a flipped
engine rather than as a too-small fixture: floor-at-the-bottom WITH gaps is a
broken premise, floor-at-the-bottom with NO gaps is a vacuous pass that must
fail loudly and say what to widen.

RESIDUAL, recorded at the call site: a BYO customer already on a future
flipped-default engine is unprotected at RUNTIME -- the pin catches it in CI,
not on their box. Only version-gated explicit options would close that,
deliberately not built because the compatibility cost is real and the trigger
speculative.

Full solution rebuild 0 warnings / 0 errors; Darling suite 3737 passed / 0
failed against live PostgreSQL 18.4 + TimescaleDB 2.28.1; the new pin run
three times for flake.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
verb-reviewer's finding, and it was right: the disclosure half never existed.
onNewestFirstUnavailable was a TRAILING OPTIONAL Action<string>?, both
production call sites simply omitted it, the null-conditional invoke was a
no-op, and a pre-2.21 store degraded in complete silence -- while the call-site
doc and a passing pin both asserted the contract WAS carried. That is worse
than not degrading at all: it is my own sentence about telling an operator a
contract is held when it is not, reappearing in the code that sentence was
written to justify.

Restores the explicit option (I had reverted it wholesale on the earlier
ruling) and fixes what neither round had caught: the capability is now LATCHED
per run in RefreshDisclosure, so a pre-2.21 store pays ONE failed call rather
than one per slice -- my previous implementation re-attempted the 5-argument
form on every slice of a potentially multi-hundred-slice backfill.

The sink is a REQUIRED constructor argument and a REQUIRED parameter on both
refresh entry points. A trailing optional is exactly the shape that let this
past a green suite, a green CI run and a self-review, because nothing anywhere
had to acknowledge it existed. Required, the compiler forces every call site --
tests included -- to decide where the disclosure goes. Enforce with the
compiler, not with a test someone has to remember to write.

Tests updated to match, and the test sink is not a stub: SilentDisclosure()
FAILS if it ever fires, because on the bundled 2.28.1 these tests must take
the explicit-options path -- a disclosure there would mean every "the option
is carried" claim in the file is hollow.

New pin covers the seam the SQL-shape pins could not: the latch flips, and the
operator is told exactly ONCE however many slices and rollups follow (a
365-slice backfill must not print 365 identical notes), plus a null sink
throws rather than silently accepting no disclosure.

The behavioural mid-slice pin STAYS despite the reviewer withdrawing it for
the declared path. It now guards the DEGRADED path, which is the one where no
option is passed at all -- BYO stores below 2.21, exactly where a mechanism
pin governs nothing.

Watched red: omitting the argument at a verb call site is now a COMPILE error
(CS7036), which is the enforcement itself rather than a test of it.

Full solution rebuild 0 warnings / 0 errors; Darling suite 3738 passed / 0
failed against live PostgreSQL 18.4 + TimescaleDB 2.28.1.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ment

Reverts ad4bea7 in full. verb-reviewer ACKED 3c3907f -- the NO-OPTIONS
state -- and explicitly verified the revert was clean and that its
dead-disclosure finding was resolved by REMOVAL rather than by wiring. I had
since restored the explicit option and wired the disclosure, which is a
substantive change to the thing that was acked, so arming on that ack would
have been arming something nobody reviewed.

On the merits I now think the reviewer and the coordinator are right and my
restore was wrong. My argument for the explicit option was runtime protection
against a default flip. The reviewer then proved by EXECUTION that the
behavioural pin catches exactly that -- flipping the engine with
refresh_newest_first=false, which 2.28.1 accepts, so not a simulation of the
threat but the threat itself -- and the pin went red with the correct
diagnosis. With the guard demonstrated against the real failure, the option's
marginal value no longer justifies a 2.21 floor on bring-your-own stores.
The tree is byte-identical to the acked SHA.

The reviewer's non-blocking poll-condition note is recorded as a comment
rather than taken as code, which it offered as an acceptable outcome. I tried
the one-liner: breaking on `floorDuring is not null` alone cancels BEFORE the
first batch commits, so a flipped engine then reports "materialized nothing at
all" -- a worse diagnosis than the fixture-size one it replaces. Doing it
properly needs a wait-for-first-commit signal that is not the floor itself.
The comment records the limitation, both measured diagnoses (120 days ->
"widen", 700 -> the true cause), and why the obvious fix is not one, so the
next person does not re-derive it.

Full solution rebuild 0 warnings / 0 errors; Darling suite 3737 passed / 0
failed, matching the reviewer's independently measured baseline exactly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Test-only. The darling-pg leg failed once on e7576d9 with two OTHER lanes'
shared-store tests ("Connection is not open", and a count reading 0 instead of
9); a re-run of the same SHA passed, and my production code there is
byte-identical to the acked 3c3907f whose CI was green -- so this is the
connection-level flake class on this rig, not a regression.

Not dismissing it on the green retry, per the rule that cost me a real hazard
earlier tonight: what did I introduce that could plausibly contribute? The
mid-slice pin deliberately ABORTS a statement mid-flight, and `await refresh`
returns when the client-side task completes -- which is not the instant the
server-side backend finishes unwinding an aborted CALL. ScratchPostgres then
ends the test with DROP DATABASE ... WITH (FORCE). Disposal order happened to
close that window; now it is closed explicitly instead of by luck.

Same lifecycle-hardening the arming tests got, and for the same reason: a test
that deliberately aborts a statement should not leave its cleanup to chance on
a rig whose flake class is connection-level. Still not a retry wrapper.

Full solution rebuild 0 warnings / 0 errors; Darling suite 3737 passed / 0
failed, matching the reviewer's independently measured baseline.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@erikdarlingdata
erikdarlingdata merged commit a0d83d7 into dev Jul 28, 2026
4 checks passed
@erikdarlingdata
erikdarlingdata deleted the feature/1759-coverage-routing-and-backfill branch July 28, 2026 07:02
erikdarlingdata added a commit that referenced this pull request Jul 28, 2026
…er-count

Correct the #1788 reader count in the CHANGELOG: six, not seven
pull Bot pushed a commit to ehtick/PerformanceMonitor that referenced this pull request Jul 29, 2026
Reapplies ad4bea7, which the coordinator's arbitration ACCEPTED and
verb-reviewer then technically acked by execution — but which was NOT on dev,
because I armed erikdarlingdata#1788 before either verdict arrived and merged the reverted
shape (b0392f7) instead. Nothing on dev is wrong; it is the shape the reviewer
acked at 3c3907f and verified red by flipping the engine. What was missing is
the runtime half.

Mechanically a revert of e7576d9, so the delta is exactly the accepted one,
plus the connection hardening that landed on dev after it.

What it restores, and why each part is not optional:

- refresh_newest_first is DECLARED, not inherited. The resume story depends on
  a killed slice leaving its NEWEST batches committed; that is not a GUC, it
  rides the options jsonb, so omitting the parameter means trusting an
  undeclared engine default. A declared option can only be broken by an engine
  ignoring its own documented contract -- an undeclared default can flip in a
  release note.
- The capability is LATCHED PER RUN. This is what dissolved the compatibility
  objection the earlier revert was ruled on: my first cut re-attempted the
  5-argument form on every slice, so a 2.18-2.20 store would have paid a failed
  round trip per slice. verb-reviewer measured the fixed shape on a simulated
  pre-2.21 store: 152 slices, ONE 42883, ONE disclosure line, DONE with zero
  missing buckets. The degraded path does not merely exist, it converges.
- The disclosure sink is REQUIRED by the compiler. A trailing optional is an
  invisible decision: the first cut made it optional, both call sites omitted
  it, and a green suite plus green CI plus a self-review all missed that the
  degrade was silent. Omitting it is now CS7036 -- the enforcement is the
  watched-red, not a test someone must remember.
- The test sink FAILS IF IT FIRES, so tests on 2.28.1 cannot quietly take the
  degraded path while asserting the option is carried.
- The behavioural mid-slice pin STAYS. verb-reviewer withdrew its withdrawal:
  on the declared path the pin is redundant, but on the DEGRADED path no option
  is passed at all, so it is the one configuration where the premise is still
  inherited -- and it is now demonstrably reachable, not theoretical.

Full solution rebuild 0 warnings / 0 errors. Darling suite: all 36 rollup
backfill tests pass. One unrelated live test (PayloadDimensionLiveTests
DimensionGc_Defers...) fails on this local rig at connection-open, BEFORE any
code under test runs -- verified pre-existing by running it on plain dev with
this change stashed, where it fails identically. CI is the authority on it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
pull Bot pushed a commit to ehtick/PerformanceMonitor that referenced this pull request Jul 29, 2026
…not seven

The shipped entry claims "All seven production routing readers are gated".
There are six. Review caught the miscount while erikdarlingdata#1788 was open and the PR body
was corrected, but the CHANGELOG kept the stale number and merged with it.

Six is what the code says, three independent ways: `coverage.For(` appears at
exactly six production call sites (ComposeSourceRouter, the MCP
DarlingHealthReader, the viewer's DailySummary, and three in
FinOps.Workload); the guard test's rot floor is pinned at `scanned >= 6`; and
that guard's own comment enumerates them as "the composer, the MCP
daily-health reader, the viewer's calendar, and three FinOps readers".

CHANGELOG-only, and deliberately no separate entry for this PR — the fix IS
the entry, and a release note announcing a corrected numeral in an unreleased
note is noise in the copy Erik publishes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
pull Bot pushed a commit to ehtick/PerformanceMonitor that referenced this pull request Jul 29, 2026
…rikdarlingdata#1795)

The erikdarlingdata#1782 guard deferred the whole dimension GC whenever a dim-feeding
purge failed - and the erikdarlingdata#1784 coverage clamp holds those purges EVERY
sweep on a coverage-lagging store, so the GC deferred every sweep and a
400-day orphan survived with nothing failed anywhere.

The GC now measures the true safety boundary instead of assuming one:
min(collection_time) per dim-feeding table under exactly the predicate
of a new V39 partial index (index-edge probe, once per sweep), minimum
across tables, clamping the assumed cutoff to one day before it. Held
history bounds the GC instead of stopping it; referenced content
survives; orphans reclaim. An UNMEASURABLE floor (table missing) still
defers - pruning on an unknown boundary is how digests dangle.

Viewer schema ladder gains the V39 arm (index-existence sentinel).
Probe predicate, index predicate, and dimension map pinned three ways.

Test-fixture defect fixed en route: EnsureContinuousAggregates attaches
refresh policies whose jobs fire immediately; the class's force-refresh
collided (55P03) and a restore DROP could strand db-grain aggregates in
the shared fixture, flipping the erikdarlingdata#1784 gate for later tests. The ensure
wrapper now removes rollup policies (tests refresh manually) and the
force-refresh retries bounded on 55P03 only - the product's erikdarlingdata#1788 idiom.

Verified: new live test proves orphan-pruned + referenced-kept while the
clamp holds; rewritten deferral test proves the unmeasurable-floor path;
both watched RED by mutating the cutoff to ignore the floor; 3x
consecutive full-class live runs, zero stranded aggregates; full fast
suite 3569 green; service + viewer builds zero warnings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
erikdarlingdata added a commit that referenced this pull request Jul 31, 2026
… stop racing it

Pre-existing flake on dev, caught by CI on this branch and root-caused here.
The same test failed on dev at 7e40c85 (an ancestor of this branch's base)
with Expected 2 / Actual 3, and on this branch with Expected 1 / Actual 0 on
the very next line. One cause, two failures that look nothing like each other.

add_compression_policy creates its job SCHEDULED with no initial_start, and
TimescaleDB launches it within a second or two - the #1788 behaviour. All three
of these tests added the policy AFTER inserting eligible chunks, so that
background run had chunks to compress and competed with the deterministic
foreground run_job the tests are built around. The background session carries
no lock_timeout (the default is wait-forever), so in the isolation test it
queued behind the ACCESS EXCLUSIVE lock the test takes on the middle chunk and
compressed it the instant the test rolled its blocker back, landing directly on
the assertions: 3 compressed if it beat the first one, a torn 2-then-0 if the
chunk flipped between the two reads.

Reproduced by replaying the test's exact two-session sequence in SQL. The
pre-fix ordering settles at 3 compressed / 0 uncompressed - dev's failure
exactly. With the policy added and parked before any chunk exists it holds a
stable 2 compressed / 1 uncompressed across every sample.

The three tests that insert chunks now add the policy first, while there is
nothing to compress, and park its job (scheduled => false, next_start =>
'infinity') before any row is inserted. The ordering is the load-bearing half:
parking a job that has already launched does not recall the run in flight. Same
idiom the file already used for the #1760 sentinel probe, and the same lever
PayloadDimensionLiveTests.EnsureAggregatesWithoutPoliciesAsync pulls against
this behaviour on the continuous-aggregate side. run_job still executes a
parked job, verified live, so the foreground path is unchanged.

It never reproduced locally because the test cluster runs max_worker_processes
= 8 against TimescaleDB's default max_background_workers = 16, so job launches
routinely fail outright. Twenty consecutive local runs passed for that reason
alone; raising the limit on the same rig made the background run fire every
time. That provisioning gap is filed as #1888.

Closes #1889.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
ianwalkeruk pushed a commit to ianwalkeruk/PerformanceMonitor that referenced this pull request Jul 31, 2026
…a#1889's fourth site (erikdarlingdata#1888 follow-through)

StuckCompressionJobsSql_NeverRunJob pins the erikdarlingdata#1760 sentinel: job_stats reads
last_run_started_at = '-infinity', not NULL, for a job that has never executed.
It created its policy UNPARKED, and add_compression_policy creates the job
SCHEDULED with no initial_start, so TimescaleDB launches it within a second or
two (erikdarlingdata#1788). One background launch destroys the sentinel outright - unlike a
chunk count, no amount of re-reading recovers it.

erikdarlingdata#1889 fixed exactly this for three sibling tests by creating and parking in ONE
transaction, so the scheduler (a separate backend) cannot see the job until the
row already reads scheduled = false. This was the fourth site and now uses the
same helper.

erikdarlingdata#1888 predicted it when it raised CI's worker slots: "any other test that has
been quietly relying on the scheduler being unable to run is going to start
failing ... whether others exist is unknown, and finding out is the actual work
here." It surfaced on a batch-two full-suite run, in a file this batch does not
otherwise touch.

Proven both ways rather than by re-running until green: with the old unparked
creation plus a deliberate four-second scheduler window the assertion fails as
Expected: True / Actual: False, byte-identical to the intermittent full-suite
failure; with the parked creation and the same window it passes 3/3.

No permanent guard, for cause: three other live tests in this file create
unparked policies BY DESIGN - an already-armed legacy policy is what the
converge tests converge - so a source-level rule would either fire on those or
be narrowed until it pinned nothing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant