Repository navigation
Rollups serve only what they materialized: route by coverage, and back fill behind a disk preflight - #1788
Merged
erikdarlingdata merged 17 commits intoJul 28, 2026
Conversation
A continuous aggregate created WITH NO DATA over pre-existing history serves only what it materialized. Real-time aggregation cannot rescue the rest: the watermark is a hard partition (materialized below UNION ALL raw at-or-above), so history below it that was never materialized is served by NEITHER branch. Every rollup's refresh policy starts 3 days back, so on a store that existed before its rollups the materialized span begins at creation-minus-3-days and never reaches further back on its own. Age-only routing therefore sent old windows to a rollup that answered with silence while raw still held every row -- and raw still held them precisely because the #1680 arming gate had held its purge closed for the same reason. RollupCoverage probes each rollup's min materialized bucket plus each rolled raw table's oldest row, and RetentionTierRouter degrades a window to the tier measured to reach furthest back. The hard half is NOT falling back when raw would return LESS. On a healthy store raw keeps ~4 days against the rollups' weeks, so a window older than every floor is the NORMAL case there and a naive floor check would send every long window to a 4-day table. The rule is comparative: a tier is abandoned only on a positive measurement that a lower tier is deeper. Unknown coverage (failed probe, plain PostgreSQL, partial build) is inert by construction. Also: - All 7 production routing readers gated, including the MCP get_daily_health reader, which answers the same question as the viewer's calendar off the same SQL -- gating one and not the other would have them disagree about how many queries ran on a given day, on exactly the affected stores. A source-parsing guard keeps a future reader from being added un-gated. - Coverage expires on the existing 5-minute reprobe interval even when availability does not. Availability is permanent once complete (a created aggregate is never dropped); coverage MOVES -- backwards on a backfill, forwards on a retention drop -- so the AllPresent permanent-cache shortcut would have pinned a pre-backfill floor for the life of the process. - The partial-window notice now prefers the MEASURED floor over the retention span. On these stores the purges are held, so raw holds months rather than its nominal ~4 days, and the assumed span would have stamped "older points are not included" on precisely the fallback panels that are complete. - Retires the false premise at the CAGG definitions ("real-time aggregation opted into... correct to query for any window immediately") and the cold-start caveat prescribing a manual backfill no store ever received. The materialized_only test pin stays, now load-bearing and for the true reason: the coverage probe and the arming gate both read min(bucket) to mean "the oldest bucket MATERIALIZED", which the real-time union branch would break. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
With reads already correct (Phase 1), materializing the rollups becomes a capacity operation rather than a correctness emergency -- which is what lets it be preflighted and operator-triggered instead of run at startup. It is NOT a startup step, and that is the whole shape of it. The #1680 arming gate is all-or-nothing, so a store with a year of raw has to materialize the WHOLE history before the first purge arms and reclaims anything: peak disk comes BEFORE any relief. Doing that automatically at service start, on the exact stores worst affected (one already down to ~150 GB free), is a plausible disk-exhaustion event. - DISK PREFLIGHT that refuses with numbers. The estimate is CALIBRATED from what the rollup has already materialized (bytes per bucket x buckets to add), which every affected store has, since its refresh policy has been materializing a trailing 3-day window all along; with no sample it bounds from raw's own size and flags itself as a bound. Requires the estimate plus 25% headroom plus a 10 GB reserve. Free space is measured on the volume the STORE says it lives on (current_setting('data_directory')), so a store on another host refuses rather than measuring this machine's disk. The refusal names the shortfall and both real options, including that waiting is SAFE -- nothing is being lost while the purges are held. - SLICED one source chunk at a time, oldest first. The engine already batches within a call, so this is supervision, not re-implementation: progress on a multi-hour run, a resume point, a lock window short enough not to sit across a compression job (#1778), and no unbounded transaction. - CONVERGENCE READ FROM DATA. A refresh that stops on its internal batch cap logs server-side and returns success to the client, so the calls returning is no evidence. Only a measured shortfall escalates to the forced form, which is the one thing that repairs a hole left by an interrupted pass. - IDEMPOTENT and resumable: every pass re-plans from the measured floor. - ARMS NOTHING. Coverage is the whole job; the arming gate already self-heals on the next service start, and duplicating that decision in a second place is how the one thing protecting this data stops being the one thing. Three findings from the gated live leg (PostgreSQL 18.4 / TimescaleDB 2.28.1), none of which were visible by reading: 1. 55P03 concurrent refresh. The verb runs while the service is UP, and the aggregate's own refresh policy lands on top of a slice. Retried, bounded, transient-only -- every other SQLSTATE still fails fast. 2. 22023 "refresh window too small". A day-wide slice's ragged tail is narrower than one bucket for a DAILY rollup, and the range end is a coverage floor or "now", so it lands mid-bucket most of the time. It aborted the whole daily tier. The planned range now closes on a bucket boundary as well as opening on one. 3. The per-slice "did the floor reach this slice's start?" check was measuring the wrong thing -- a slice's range can legitimately hold no source rows (raw's oldest row lands partway into the first slice; a collection gap does the same mid-run), so it fired on the first slice of every run. A global min(bucket) cannot answer a question about one slice. Convergence is judged once, at the end, against raw's oldest row. Tested with the verb driven end to end against a real store through a bring-your-own darling.json, not a re-implementation of its loop: dry run changes nothing, the real run's DONE claim is re-verified against the store, and a second run reports nothing to do. Every guard verified RED by mutation: preflight bypassed, convergence made call-based over a silently under-covering plan, and the hierarchical order inverted (caught by the pin AND by the live verb, which reported the dailies genuinely short). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This was referenced Jul 28, 2026
…ntee Comment-only. Names the two shapes a one-dimensional floor cannot see, and blocks the inference that would let them be forgotten again. A mid-window HOLE is invisible: coverage is one number, so it cannot see a gap above itself. A service down longer than the refresh policy's 3-day start_offset resumes at now-3d and never backfills the skipped interval, while min(bucket) goes on reporting the original deep floor -- so the window routes to the tier and is served as complete with the gap inside it. A window STRADDLING the floor is served partially with no signal. Returning the tier is the correct CHOICE there (raw would return less on a healthy store), but the part below the floor is missing and nothing says so. Both are accepted: both are strictly better than the age-only routing they replace, which served the ENTIRE window as empty in exactly these cases. What the paragraph exists to block is the belief "the coverage gate passed, therefore the result is complete" -- which is precisely how #1759 survived as long as it did. The previous comment asserted the view was "correct to query for any window immediately", it read as reasonable, it was pinned by a test, and nobody re-derived it for months. A one-dimensional measure invites the same inference in a new place. Both gaps and their candidate closures are tracked in #1791 rather than assumed away here. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The guard standing between #1759 and its own return could be satisfied by the most probable way back. It checked that a Resolve call's argument list contained the substring "overage" -- which is a substring of the TYPE NAME TierCoverage. So a reader passing TierCoverage.Unknown, the canonical "route on no evidence" value, SATISFIED the guard. That form compiles and reads as deliberate, which is exactly what makes it likely: not someone dropping an argument (the form this had been verified against, and one a compile-shaped review also catches), but someone reaching for the inert value because it satisfied the signature. The predicate is now two-sided -- a positive match on a real coverage lookup (.For() AND an explicit rejection of TierCoverage.Unknown. Either half alone is holed: "mentions coverage" is satisfied by the type name, and "does not say Unknown" is satisfied by passing nothing at all. The failure message names which of the two it found, since they need different fixes. Tightening it surfaced a second flaw the one-sided predicate had been hiding: the scan matched <see cref="...Resolve(...)"/> inside DOC COMMENTS and, since the pattern spans newlines looking for the closing ");", ran that phantom match forward into whatever real code followed. It passed silently because a cref contains the type name -- and it inflated the rot detector's count. Comments are stripped before scanning now, preserving line numbers so reports still point somewhere openable. Rot detector raised 5 -> 6, the verified production call-site count: the composer (one site, routing dynamically across all three catalog tables), the MCP daily-health reader, the viewer's calendar, and three FinOps readers. A legitimate removal should lower it deliberately rather than be absorbed. Both forms verified RED against a green baseline, each naming its own cause. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Comment-only. #1776 landed a hygiene guard on dev while this branch was open: every class reaching the shared DARLING_TEST_PG store must either serialize with [Collection("live-postgres")] or say in writing why it does not. This class mints its own scratch database, so it cannot race the shared store and serializing it would be pure slowdown -- it is the exemption case, and it now carries the "#1776 own-store" marker the guard looks for. It has to own its store regardless: it creates continuous aggregates and retention policies the shared fixture must never inherit from a test. Found by test-merging origin/dev rather than trusting the merge status. The merge is TEXTUALLY clean on everything but a keep-both CHANGELOG link-ref block -- git reported no conflict here at all, because a new guard on one side and a new class on the other do not overlap as text. Only building and running the merged tree surfaced it. Fixing it here rather than at merge time keeps the eventual resolve to the CHANGELOG alone. Verified on the merged tree: LivePostgresCollectionHygieneTests passes and the full suite is green (3546 passed, 0 failed). This branch alone: 3537 passed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Resolved to get CI to run at all, not as an arm-time step: a CONFLICTING PR has no merge ref, and pull_request workflows build the merge ref -- so Build and Claude Auto Review could not start on this PR while the conflict stood, while other branches' runs kept firing normally. That is what made this look like an Actions outage for ~25 minutes. CHANGELOG.md was the only conflict, and only in the link-ref block. Resolved KEEP-BOTH with no re-sorting: this branch's #1665/#1788 refs and dev's #1781/#1783/#1786/#1792 refs all retained, in the order they appeared. No entry text on either side was touched. Everything else auto-merged. Verified rather than assumed, because the dangerous case here is textual cleanliness hiding a semantic break: full solution rebuild 0 warnings / 0 errors, Darling suite 3549 passed / 0 failed. The one real collision this class of merge produced -- #1776's new live-store hygiene guard versus this branch's new own-store live test -- was found the same way and fixed in 59f877a before this merge. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ill hole BLOCKING 1 -- the preflight under-estimated by rows-per-bucket. The probe counted ROWS (count(*)) and the arithmetic divided that into bytes to get a per-BUCKET figure, then multiplied by a BUCKET count. A rollup holds one row per (server, database, query_hash, sql_handle, bucket), so the two units differ by the distinct queries seen per hour -- measured 205x under on a 200-row/hour store, worse on a real fleet box. Under-estimating is the single direction this preflight exists to prevent, and my own comment said so. count(DISTINCT bucket) fixes it. The suite could not have caught it: BOTH live fixtures seeded one row per bucket, making rows and buckets the same number and the error factor exactly 1. That is not a small fixture, it is a fixture of a shape the product never meets. They now seed 12 distinct queries per bucket, and a live test asserts the probe's count is a bucket count by ratio rather than by string. BLOCKING 2 -- an interrupted backfill left a hole, reported DONE, and the arming gate then armed a purge over it. min(bucket) is the resume point, the completion verdict, AND what the #1680 gate reads before letting the 4-day raw purge drop chunks. Slices ran OLDEST-FIRST, so the FIRST slice drove that one number to its final value and every later slice was invisible to all three. Proven by execution: killed at slice 32/43 -> 47/264 buckets missing -> re-run printed "nothing to do -- DONE" and exited 0 -> the gate armed over a window whose only surviving copy was the raw rows about to be dropped. The slice-failure path reached the same state. My "Safe to interrupt" line was false. Fixed by REVERSING the slice order rather than adding per-slice bookkeeping. Descending makes floor <= raw_oldest imply completeness BY CONSTRUCTION, in any interleaving: the floor can only reach the bottom once the last, oldest slice has run. An interrupted run leaves it visibly short, which is a truthful SHORT rather than a false DONE, and the gate stays closed because it reads the same honest number. Mid-slice kills need no extra state: refresh commits per batch newest-first, so the floor lands inside the killed slice and re-planning from it re-covers the remainder. Chose this over per-slice verification because it removes the failure mode instead of detecting it -- there is no interleaving left in which a hole can be reported as coverage, so nothing depends on a probe being remembered. Also folded in the sanity clamp: one stray epoch-era row produced a 739,825-slice plan that burned CPU indefinitely. Plans past a 10-year ceiling now REFUSE and name the offending timestamp -- corruption someone must look at, not history to back fill. A refusal is deliberately not a skip: it prints [REFUSED] to stderr, blocks DONE, and exits non-zero, because burying it in the [OK] lines is how it would be missed. Every fix verified RED by mutation against a green baseline: - count(*) restored -> unit pin AND the live ratio guard both fail - ascending order restored -> the live acceptance reproduces the exact defect ("an interrupted backfill reports coverage ... at or below raw's oldest -- it looks COMPLETE"), plus two unit pins New coverage: the reviewer's proof inverted (kill at 2/3, assert the partial run does not look complete, the re-plan still has work, the REAL arming gate refuses, then resume and assert ZERO missing buckets measured against raw itself, then the gate arms), and the verb's refusal branch end to end, which previously had only its pure arithmetic covered. Full solution rebuild 0 warnings / 0 errors; Darling suite 3723 passed / 0 failed against live PostgreSQL 18.4 + TimescaleDB 2.28.1; backfill tests run three times for flake. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Resolved promptly rather than at arm time: a DIRTY PR gets no pull_request CI at all (the merge ref is uncomputable), so leaving it dirty would strand the two BLOCKING fixes in be0d2c6 without check evidence for either review seat. CHANGELOG.md auto-merged this time -- no conflict, nothing to hand-resolve. Adopted proof run and clean: `git diff origin/dev -- CHANGELOG.md` shows only the two link-refs this branch ADDS (#1665, #1788), with no line removed and none moved. Brought in from dev: DarlingRetention.cs and PayloadDimensionLiveTests.cs, neither of which this branch touches. Verified rather than trusted, because MERGEABLE is not "the merge is safe" -- a guard landing on one side and a new class on the other do not overlap as text but do collide semantically, which is exactly how #1776's hygiene guard caught this branch's own live test earlier tonight. Full solution rebuild 0 warnings / 0 errors; Darling suite 3724 passed / 0 failed against live PostgreSQL 18.4 + TimescaleDB 2.28.1, so the six #1759 live legs ran rather than skipping. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Both are properties the fix DEPENDS on that nothing was enforcing, raised by verb-reviewer's re-verification. 1. refresh_newest_first. The no-extra-state mid-slice resume argument rests on a killed slice leaving its NEWEST batches committed, so the floor lands inside that slice and the next run's top slice re-covers the remainder. That is true only because refresh_newest_first is TRUE -- which is the 2.28 DEFAULT taken when options is NULL. Verified the verb passes no 5th argument on either form, and it now says so at the call site rather than leaving the dependency invisible. The pin fails the moment anything starts passing options, at which point the assumption has to be made explicit (pass refresh_newest_first by name) instead of inherited from a value somebody else chose. Asserted on the forced form too: it already takes a 4th argument and is the likeliest place a 5th gets appended. 2. Bucket alignment across the reversal (P2b). The alignment property was established against ASCENDING slicing, where the ragged remainder sat at the TOP of the range. Reversing moved it to the BOTTOM -- a different slice, produced by a different branch of the loop (the clamp to `from` rather than the clamp to `to`) -- so the property had to be re-established rather than assumed to have carried. Now checks BOTH ENDS of every slice against the bucket grid, not just the width: a slice one bucket wide but half a bucket out of phase is still refused with 22023. Plus a pin that the remainder really is at the bottom, which would catch a silent revert to ascending putting it back on top and re-opening the false-DONE path. Both verified RED against a green baseline, on mutations that COMPILE: appending `false, $3::jsonb` to the refresh call, and restoring ascending order. The first attempt at the options mutation did not compile and the test ran against a stale binary reporting green -- worth naming, because a mutation that fails to build is indistinguishable from a guard that works. Full solution rebuild 0 warnings / 0 errors; Darling suite 3725 passed / 0 failed against live PostgreSQL 18.4 + TimescaleDB 2.28.1. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The reviewer's premise check was right and the comment was wrong. It said
refresh_newest_first "is on" as though it were configured -- but there is no
such GUC on 2.28.1 (not in pg_settings). It rides the options jsonb parameter,
which this verb never passed, so the mid-slice resume guarantee rested on an
UNSTATED SERVER DEFAULT.
It holds today, and a real mid-slice kill measured zero holes above the floor.
That is exactly what makes it worth fixing rather than leaving: this PR's
original defect came from the same class of trust. materialized_only was also
assumed, also read as reasonable, also pinned by a test, and was wrong for
months. An unstated default that happens to be right is not a contract.
So the call now passes options => {"refresh_newest_first": true} explicitly,
force becomes positional to reach it, and the comment states what the code
ASSERTS rather than what it hopes. Only newest-first is set: batch size and
the batch cap stay at the defaults the issue's API research established are
already the right shape.
Degrades honestly rather than hard-failing. options arrived in 2.21 and
nothing gates a bring-your-own store's version, so a pre-2.21 store raises
42883 on the 5-argument call. RefreshAsync falls back to the 3-argument form
ONCE and says so out loud -- on such a store the guarantee genuinely IS an
inherited default, and telling an operator a contract is held when it is not
would be the same failure in a new place. Branched on SQLSTATE, never the
message.
The pin is inverted accordingly: it previously asserted no options were
passed, and its own doc said "anything that starts passing options must pass
refresh_newest_first explicitly and turn this into a pinned property rather
than an assumption". That is what happened, so it now asserts the contract IS
carried, on both the plain and forced forms, plus the degraded 3-argument
shape and the SQLSTATE it keys on.
Verified against the real store rather than assumed: the full interrupted-
backfill acceptance re-ran on live PostgreSQL 18.4 / TimescaleDB 2.28.1 with
the explicit option in place -- 2.28.1 accepts it and the kill/resume/zero-
missing-buckets/gate-arms chain still holds. Mutation (withOptions default
flipped back to false) verified RED on a COMPILING build.
Full solution rebuild 0 warnings / 0 errors; Darling suite 3725 passed / 0
failed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Test-only hardening, prompted by a CI flake I could not honestly call unrelated. The darling-pg leg failed once on 28987fd with 15 tests dying inside one second on "existing connection was forcibly closed" -- a server-level event, not a logic failure, and mostly in shared-store classes this branch does not touch. A re-run of the same SHA passed, so it is not deterministic and the explicit-options change did not cause it. But a green retry proves only that it is not deterministic; it does not make my tests innocent, and one of the casualties WAS mine. What I did introduce is a genuinely destabilising pattern: these tests ARM real retention policies, because whether the #1680 gate DECIDES to arm is the thing under test. An armed policy means TimescaleDB's scheduler will launch a background worker to run drop_chunks -- against a scratch database ScratchPostgres is about to DROP ... WITH (FORCE). Leaving a worker running against a vanishing database is a hazard whether or not it caused this particular flake. So both arming tests now unschedule their jobs once the assertions have read the catalog. That costs the tests nothing: what is under test is the gate's DECISION, already measured by then, and nothing needs the purge to actually execute. Best-effort, because a cleanup failure that fails a passing test inverts the signal. Deliberately NOT a retry wrapper. The recorded lesson from the earlier live-PG flake is to harden the class's lifecycle rather than paper over it, and a retry here would have hidden exactly the pattern worth removing. Full solution rebuild 0 warnings / 0 errors; Darling suite 3725 passed / 0 failed against live PostgreSQL 18.4 + TimescaleDB 2.28.1. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Resolve-only push during an active review window. CHANGELOG auto-merged with no conflict; adopted proof clean (only this branch's #1665/#1788 link-refs added, none removed or moved). Dev brought in the dims pair plus TimescaleSupport.cs, which this branch also edits -- so verified rather than trusted, per MERGEABLE-is-not-merge-safe: full solution rebuild 0 warnings / 0 errors, Darling suite 3736 passed / 0 failed against live PostgreSQL 18.4 + TimescaleDB 2.28.1 with the live legs actually running. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two decisions, both reversals of my own previous commit. REVERTED the explicit options argument. The compatibility matrix corrected both the prescription and my implementation of it: options arrived in TimescaleDB 2.21, nothing in this product gates a bring-your-own store's version, and my degrade-on-42883 fallback re-attempted the 5-argument form on EVERY slice -- so a 2.18-2.20 store would have eaten a failed round trip per slice, for a backfill that already works there. That is a real cost paid to guard an engine flip that has not happened. The no-options shape is restored; the comment stays corrected, because the original lie was calling an inherited ENGINE DEFAULT a configured setting, and that is still worth saying plainly. REPLACED it with a guard on the BEHAVIOUR. Every other test in this file kills BETWEEN slices, so the mid-slice premise -- a kill inside a slice leaves the floor inside the cancelled range with no holes above it -- had nothing watching it; its only evidence was hand-executed kills in a transcript. The new pin cancels a wide refresh mid-flight and asserts exactly that. It never mentions refresh_newest_first, so it holds on any engine including BYO stores that would reject the option outright, and a bundled-runtime bump re-runs it automatically. THE FIXTURE IS THE WHOLE TEST, and my first two attempts were vacuous. The ensure sweep also attaches a refresh policy, TimescaleDB runs a new policy's first check IMMEDIATELY, and that policy materializes a trailing window CONTIGUOUSLY -- which satisfies both assertions on its own, whatever the cancelled refresh did. An oldest-first mutation PASSED twice: once with the policy live, and again after unscheduling it post-creation, which is too late. The aggregate is now created from its own DDL with NO policy attached, so the only materialization in the database is the one being cancelled. Only then did the mutation go red -- 2832 missing buckets, correctly diagnosed as "this engine did NOT commit its batches newest-first". Also reordered the assertions so a flipped engine is diagnosed as a flipped engine rather than as a too-small fixture: floor-at-the-bottom WITH gaps is a broken premise, floor-at-the-bottom with NO gaps is a vacuous pass that must fail loudly and say what to widen. RESIDUAL, recorded at the call site: a BYO customer already on a future flipped-default engine is unprotected at RUNTIME -- the pin catches it in CI, not on their box. Only version-gated explicit options would close that, deliberately not built because the compatibility cost is real and the trigger speculative. Full solution rebuild 0 warnings / 0 errors; Darling suite 3737 passed / 0 failed against live PostgreSQL 18.4 + TimescaleDB 2.28.1; the new pin run three times for flake. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
verb-reviewer's finding, and it was right: the disclosure half never existed. onNewestFirstUnavailable was a TRAILING OPTIONAL Action<string>?, both production call sites simply omitted it, the null-conditional invoke was a no-op, and a pre-2.21 store degraded in complete silence -- while the call-site doc and a passing pin both asserted the contract WAS carried. That is worse than not degrading at all: it is my own sentence about telling an operator a contract is held when it is not, reappearing in the code that sentence was written to justify. Restores the explicit option (I had reverted it wholesale on the earlier ruling) and fixes what neither round had caught: the capability is now LATCHED per run in RefreshDisclosure, so a pre-2.21 store pays ONE failed call rather than one per slice -- my previous implementation re-attempted the 5-argument form on every slice of a potentially multi-hundred-slice backfill. The sink is a REQUIRED constructor argument and a REQUIRED parameter on both refresh entry points. A trailing optional is exactly the shape that let this past a green suite, a green CI run and a self-review, because nothing anywhere had to acknowledge it existed. Required, the compiler forces every call site -- tests included -- to decide where the disclosure goes. Enforce with the compiler, not with a test someone has to remember to write. Tests updated to match, and the test sink is not a stub: SilentDisclosure() FAILS if it ever fires, because on the bundled 2.28.1 these tests must take the explicit-options path -- a disclosure there would mean every "the option is carried" claim in the file is hollow. New pin covers the seam the SQL-shape pins could not: the latch flips, and the operator is told exactly ONCE however many slices and rollups follow (a 365-slice backfill must not print 365 identical notes), plus a null sink throws rather than silently accepting no disclosure. The behavioural mid-slice pin STAYS despite the reviewer withdrawing it for the declared path. It now guards the DEGRADED path, which is the one where no option is passed at all -- BYO stores below 2.21, exactly where a mechanism pin governs nothing. Watched red: omitting the argument at a verb call site is now a COMPILE error (CS7036), which is the enforcement itself rather than a test of it. Full solution rebuild 0 warnings / 0 errors; Darling suite 3738 passed / 0 failed against live PostgreSQL 18.4 + TimescaleDB 2.28.1. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ment Reverts ad4bea7 in full. verb-reviewer ACKED 3c3907f -- the NO-OPTIONS state -- and explicitly verified the revert was clean and that its dead-disclosure finding was resolved by REMOVAL rather than by wiring. I had since restored the explicit option and wired the disclosure, which is a substantive change to the thing that was acked, so arming on that ack would have been arming something nobody reviewed. On the merits I now think the reviewer and the coordinator are right and my restore was wrong. My argument for the explicit option was runtime protection against a default flip. The reviewer then proved by EXECUTION that the behavioural pin catches exactly that -- flipping the engine with refresh_newest_first=false, which 2.28.1 accepts, so not a simulation of the threat but the threat itself -- and the pin went red with the correct diagnosis. With the guard demonstrated against the real failure, the option's marginal value no longer justifies a 2.21 floor on bring-your-own stores. The tree is byte-identical to the acked SHA. The reviewer's non-blocking poll-condition note is recorded as a comment rather than taken as code, which it offered as an acceptable outcome. I tried the one-liner: breaking on `floorDuring is not null` alone cancels BEFORE the first batch commits, so a flipped engine then reports "materialized nothing at all" -- a worse diagnosis than the fixture-size one it replaces. Doing it properly needs a wait-for-first-commit signal that is not the floor itself. The comment records the limitation, both measured diagnoses (120 days -> "widen", 700 -> the true cause), and why the obvious fix is not one, so the next person does not re-derive it. Full solution rebuild 0 warnings / 0 errors; Darling suite 3737 passed / 0 failed, matching the reviewer's independently measured baseline exactly. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Test-only. The darling-pg leg failed once on e7576d9 with two OTHER lanes' shared-store tests ("Connection is not open", and a count reading 0 instead of 9); a re-run of the same SHA passed, and my production code there is byte-identical to the acked 3c3907f whose CI was green -- so this is the connection-level flake class on this rig, not a regression. Not dismissing it on the green retry, per the rule that cost me a real hazard earlier tonight: what did I introduce that could plausibly contribute? The mid-slice pin deliberately ABORTS a statement mid-flight, and `await refresh` returns when the client-side task completes -- which is not the instant the server-side backend finishes unwinding an aborted CALL. ScratchPostgres then ends the test with DROP DATABASE ... WITH (FORCE). Disposal order happened to close that window; now it is closed explicitly instead of by luck. Same lifecycle-hardening the arming tests got, and for the same reason: a test that deliberately aborts a statement should not leave its cleanup to chance on a rig whose flake class is connection-level. Still not a retry wrapper. Full solution rebuild 0 warnings / 0 errors; Darling suite 3737 passed / 0 failed, matching the reviewer's independently measured baseline. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This was referenced Jul 28, 2026
erikdarlingdata
added a commit
that referenced
this pull request
Jul 28, 2026
…er-count Correct the #1788 reader count in the CHANGELOG: six, not seven
This was referenced Jul 28, 2026
Merged
pull Bot
pushed a commit
to ehtick/PerformanceMonitor
that referenced
this pull request
Jul 29, 2026
Reapplies ad4bea7, which the coordinator's arbitration ACCEPTED and verb-reviewer then technically acked by execution — but which was NOT on dev, because I armed erikdarlingdata#1788 before either verdict arrived and merged the reverted shape (b0392f7) instead. Nothing on dev is wrong; it is the shape the reviewer acked at 3c3907f and verified red by flipping the engine. What was missing is the runtime half. Mechanically a revert of e7576d9, so the delta is exactly the accepted one, plus the connection hardening that landed on dev after it. What it restores, and why each part is not optional: - refresh_newest_first is DECLARED, not inherited. The resume story depends on a killed slice leaving its NEWEST batches committed; that is not a GUC, it rides the options jsonb, so omitting the parameter means trusting an undeclared engine default. A declared option can only be broken by an engine ignoring its own documented contract -- an undeclared default can flip in a release note. - The capability is LATCHED PER RUN. This is what dissolved the compatibility objection the earlier revert was ruled on: my first cut re-attempted the 5-argument form on every slice, so a 2.18-2.20 store would have paid a failed round trip per slice. verb-reviewer measured the fixed shape on a simulated pre-2.21 store: 152 slices, ONE 42883, ONE disclosure line, DONE with zero missing buckets. The degraded path does not merely exist, it converges. - The disclosure sink is REQUIRED by the compiler. A trailing optional is an invisible decision: the first cut made it optional, both call sites omitted it, and a green suite plus green CI plus a self-review all missed that the degrade was silent. Omitting it is now CS7036 -- the enforcement is the watched-red, not a test someone must remember. - The test sink FAILS IF IT FIRES, so tests on 2.28.1 cannot quietly take the degraded path while asserting the option is carried. - The behavioural mid-slice pin STAYS. verb-reviewer withdrew its withdrawal: on the declared path the pin is redundant, but on the DEGRADED path no option is passed at all, so it is the one configuration where the premise is still inherited -- and it is now demonstrably reachable, not theoretical. Full solution rebuild 0 warnings / 0 errors. Darling suite: all 36 rollup backfill tests pass. One unrelated live test (PayloadDimensionLiveTests DimensionGc_Defers...) fails on this local rig at connection-open, BEFORE any code under test runs -- verified pre-existing by running it on plain dev with this change stashed, where it fails identically. CI is the authority on it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
pull Bot
pushed a commit
to ehtick/PerformanceMonitor
that referenced
this pull request
Jul 29, 2026
…not seven The shipped entry claims "All seven production routing readers are gated". There are six. Review caught the miscount while erikdarlingdata#1788 was open and the PR body was corrected, but the CHANGELOG kept the stale number and merged with it. Six is what the code says, three independent ways: `coverage.For(` appears at exactly six production call sites (ComposeSourceRouter, the MCP DarlingHealthReader, the viewer's DailySummary, and three in FinOps.Workload); the guard test's rot floor is pinned at `scanned >= 6`; and that guard's own comment enumerates them as "the composer, the MCP daily-health reader, the viewer's calendar, and three FinOps readers". CHANGELOG-only, and deliberately no separate entry for this PR — the fix IS the entry, and a release note announcing a corrected numeral in an unreleased note is noise in the copy Erik publishes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
pull Bot
pushed a commit
to ehtick/PerformanceMonitor
that referenced
this pull request
Jul 29, 2026
…rikdarlingdata#1795) The erikdarlingdata#1782 guard deferred the whole dimension GC whenever a dim-feeding purge failed - and the erikdarlingdata#1784 coverage clamp holds those purges EVERY sweep on a coverage-lagging store, so the GC deferred every sweep and a 400-day orphan survived with nothing failed anywhere. The GC now measures the true safety boundary instead of assuming one: min(collection_time) per dim-feeding table under exactly the predicate of a new V39 partial index (index-edge probe, once per sweep), minimum across tables, clamping the assumed cutoff to one day before it. Held history bounds the GC instead of stopping it; referenced content survives; orphans reclaim. An UNMEASURABLE floor (table missing) still defers - pruning on an unknown boundary is how digests dangle. Viewer schema ladder gains the V39 arm (index-existence sentinel). Probe predicate, index predicate, and dimension map pinned three ways. Test-fixture defect fixed en route: EnsureContinuousAggregates attaches refresh policies whose jobs fire immediately; the class's force-refresh collided (55P03) and a restore DROP could strand db-grain aggregates in the shared fixture, flipping the erikdarlingdata#1784 gate for later tests. The ensure wrapper now removes rollup policies (tests refresh manually) and the force-refresh retries bounded on 55P03 only - the product's erikdarlingdata#1788 idiom. Verified: new live test proves orphan-pruned + referenced-kept while the clamp holds; rewritten deferral test proves the unmeasurable-floor path; both watched RED by mutating the cutoff to ignore the floor; 3x consecutive full-class live runs, zero stranded aggregates; full fast suite 3569 green; service + viewer builds zero warnings. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This was referenced Jul 31, 2026
Correct the Query Store rollups: dedup cumulative interval snapshots in the CAGG chain (#1849)
#1870
Merged
erikdarlingdata
added a commit
that referenced
this pull request
Jul 31, 2026
… stop racing it Pre-existing flake on dev, caught by CI on this branch and root-caused here. The same test failed on dev at 7e40c85 (an ancestor of this branch's base) with Expected 2 / Actual 3, and on this branch with Expected 1 / Actual 0 on the very next line. One cause, two failures that look nothing like each other. add_compression_policy creates its job SCHEDULED with no initial_start, and TimescaleDB launches it within a second or two - the #1788 behaviour. All three of these tests added the policy AFTER inserting eligible chunks, so that background run had chunks to compress and competed with the deterministic foreground run_job the tests are built around. The background session carries no lock_timeout (the default is wait-forever), so in the isolation test it queued behind the ACCESS EXCLUSIVE lock the test takes on the middle chunk and compressed it the instant the test rolled its blocker back, landing directly on the assertions: 3 compressed if it beat the first one, a torn 2-then-0 if the chunk flipped between the two reads. Reproduced by replaying the test's exact two-session sequence in SQL. The pre-fix ordering settles at 3 compressed / 0 uncompressed - dev's failure exactly. With the policy added and parked before any chunk exists it holds a stable 2 compressed / 1 uncompressed across every sample. The three tests that insert chunks now add the policy first, while there is nothing to compress, and park its job (scheduled => false, next_start => 'infinity') before any row is inserted. The ordering is the load-bearing half: parking a job that has already launched does not recall the run in flight. Same idiom the file already used for the #1760 sentinel probe, and the same lever PayloadDimensionLiveTests.EnsureAggregatesWithoutPoliciesAsync pulls against this behaviour on the continuous-aggregate side. run_job still executes a parked job, verified live, so the foreground path is unchanged. It never reproduced locally because the test cluster runs max_worker_processes = 8 against TimescaleDB's default max_background_workers = 16, so job launches routinely fail outright. Twenty consecutive local runs passed for that reason alone; raising the limit on the same rig made the background run fire every time. That provisioning gap is filed as #1888. Closes #1889. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This was referenced Jul 31, 2026
ianwalkeruk
pushed a commit
to ianwalkeruk/PerformanceMonitor
that referenced
this pull request
Jul 31, 2026
…a#1889's fourth site (erikdarlingdata#1888 follow-through) StuckCompressionJobsSql_NeverRunJob pins the erikdarlingdata#1760 sentinel: job_stats reads last_run_started_at = '-infinity', not NULL, for a job that has never executed. It created its policy UNPARKED, and add_compression_policy creates the job SCHEDULED with no initial_start, so TimescaleDB launches it within a second or two (erikdarlingdata#1788). One background launch destroys the sentinel outright - unlike a chunk count, no amount of re-reading recovers it. erikdarlingdata#1889 fixed exactly this for three sibling tests by creating and parking in ONE transaction, so the scheduler (a separate backend) cannot see the job until the row already reads scheduled = false. This was the fourth site and now uses the same helper. erikdarlingdata#1888 predicted it when it raised CI's worker slots: "any other test that has been quietly relying on the scheduler being unable to run is going to start failing ... whether others exist is unknown, and finding out is the actual work here." It surfaced on a batch-two full-suite run, in a file this batch does not otherwise touch. Proven both ways rather than by re-running until green: with the old unparked creation plus a deliberate four-second scheduler window the assertion fails as Expected: True / Actual: False, byte-identical to the intermittent full-suite failure; with the parked creation and the same window it passes 3/3. No permanent guard, for cause: three other live tests in this file create unparked policies BY DESIGN - an already-armed legacy policy is what the converge tests converge - so a source-level rule would either fire on those or be narrowed until it pinned nothing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #1759 (Phases 1 and 2). Phase 0 — the held-paused observability WARN lines — already shipped in #1762.
One PR, two commits, and why
Phase 2 hard-depends on Phase 1's coverage probe: the backfill's resume point, its convergence check and the router's fallback all read the same measured floor, and splitting them would have meant a stacked PR against a protected branch on the night before a release cut. The two commits are cleanly separated and reviewable in order — Phase 1 is read-side only and could ship alone; Phase 2 is an operator verb nothing invokes by accident.
Phase 1 — coverage-aware routing (read-side, zero disk risk)
RetentionTierRouterrouted by AGE alone, and nothing fell back to raw. On a store whose rollups were createdWITH NO DATAover pre-existing history, every window older than the rollup's materialized floor read empty while raw still held every row — and raw still held them precisely because the #1680 arming gate had held its purge closed for the same reason.RollupCoverage(a companion to #1665'sRollupAvailability) probes each rollup'smin(bucket)plus each rolled raw table'smin(collection_time)in one round trip, and the router degrades a window to whichever tier is measured to reach furthest back.The hard half is not falling back when raw would return LESS. On a healthy store with armed purges, raw keeps ~4 days against the rollups' weeks, so a window older than every floor is the normal case there — a naive "window predates the floor → use raw" rule would send every long window to a 4-day table. The rule is therefore comparative: a tier is abandoned only on a positive measurement that a lower tier is deeper, which is exactly the held-purge signature #1759 describes and is silent everywhere else. Half the routing truth table exists to hold that line, and there is a live-PG test for it specifically.
Unknown coverage is inert by construction — a failed probe, a plain-PostgreSQL store and a partially-built one all produce nulls, and nulls never move a window, so the pre-#1759 behaviour is reproduced exactly.
Other things this needed:
All six production routing call sites gated, enumerated:
ComposeSourceRouter.cs:178(one site, routing dynamically across all three catalog tables),Mcp/DarlingHealthReader.cs:210,ViewerDataService.DailySummary.cs:56, and three inViewerDataService.FinOps.Workload.cs(:205on the db-grain pair,:363and:429on the query-grain pair). (An earlier draft of this body said seven — a miscount, corrected. Six is the verified number; the guard's rot detector now pins that floor.) The MCPget_daily_healthreader is not incidental: it answers the same question as the viewer's Performance Calendar off the same shared SQL, so gating one and not the other would have had them report different query counts for the same day on exactly the affected stores, with no way to tell which was right.The source-parsing guard that keeps a future reader from being added un-gated is deliberately two-sided — a positive match on a real coverage lookup (
.For() and an explicit rejection ofTierCoverage.Unknown. Its first cut was one-sided ("does the argument list mention coverage?") and was holed:"overage"is a substring of the type nameTierCoverage, so a reader passingTierCoverage.Unknownsatisfied it. That form compiles and reads as deliberate, which makes it the most probable way Rollup CAGGs serve only materialized buckets: old windows read empty and raw purges stay held #1759 returns — not someone dropping an argument, which a compile-shaped review also catches, but someone reaching for the inert value because it satisfied the signature. Both forms are now pinned as mutations (table below). Tightening it also exposed that the scan was matching<see cref="…Resolve(…)"/>inside doc comments and running that phantom match forward into real code; the one-sided predicate had hidden it, because a cref contains the type name. Comments are stripped before scanning now.Coverage expires even where availability does not. Availability is cached permanently once complete, because a created aggregate is never dropped — true of existence, false of coverage, which moves backwards on a backfill and forwards on a retention drop. Keeping the
AllPresentshortcut would have pinned a pre-backfill floor for the life of the process, so an operator who had just backfilled would keep getting raw fallbacks until a restart. The existing 5-minute reprobe interval now applies unconditionally, in both the composer's per-datasource cache and the viewer's.The partial-window notice now prefers the MEASURED floor to the retention span. On these stores the purges are held, so raw holds months rather than its nominal ~4 days; left assuming, the notice would have stamped "older points are not included" on precisely the coverage-fallback panels that are in fact complete — a false alarm introduced by the fix. It falls back to the span when nothing is measured, which reproduces the old text.
The false premise is retired.
TimescaleSupport's claim that "real-time aggregation is opted into… correct to query for any window immediately" was wrong twice over — 2.13+ defaultsmaterialized_onlyto TRUE, and even ON the watermark is a hard partition, so un-materialized history is served by neither branch. Thematerialized_onlytest pin stays, now load-bearing and for the true reason:RollupCoverageProbeSqlandRetentionArmSafetySqlboth readmin(bucket)to mean "the oldest bucket MATERIALIZED", and unioning the raw branch in would make an empty materialization report raw's own oldest row as the rollup's floor — the router would believe coverage it does not have, and the arming gate would arm a purge over history nothing else holds. The cold-start caveat prescribing a manual backfill no store ever received is retired too, replaced by a pointer to the verb.Phase 2 —
--backfill-rollups, an operator verb with a disk preflightThe rejected option ("just open the gate") is not resurrected. The arming gate is all-or-nothing, so a store with a year of raw must materialize the whole history before the first purge arms and reclaims anything: peak disk comes before any relief. At service start, on the worst-affected stores, that is a plausible disk-exhaustion event.
current_setting('data_directory')), so a store on another host refuses rather than measuring this machine's disk. The refusal names the shortfall and both real options — including that waiting is safe, since nothing is being lost while the purges are held.--dry-runprints the plan, the estimate and a duration budget at the ~16 MB/s measured on this host class, and stops.Concurrency / locking.
refresh_continuous_aggregatetakes no lock that blocks writers on the source hypertable, so collection keeps running; it cannot run inside a transaction, so each slice is its own statement and partial progress survives an abort. What it does contend with is the compression policy on the same chunks, which is why slices are one chunk wide.The operator contract, verbatim
Every line the verb tells an operator is reproduced here rather than paraphrased, because these are the contract and a reviewer should be able to read them without opening the source. Both are pinned by
BackfillRollupsVerb_DryRunsThenBackfills_AndReportsCoverageItActuallyReached, so editing them breaks a test rather than silently voiding the contract.On completion — the next step, the confirmation to look for, and the one interaction the operator has to know about: the hourly rollups carry their own 21-day retention policy, already armed on these stores, which will trim the coverage a run just built when it next fires (measured cadence:
schedule_interval = 1 day, see #1790):And the sibling line on the nothing-to-do path (
DarlingCliCommands.cs:2187), which carries the same self-heal promise for the store that is already covered — including a re-run after a successful pass, which is the most likely way an operator sees this verb a second time:Two BLOCKING review defects, both proven by execution, both fixed
1. The preflight under-estimated by rows-per-bucket. The probe counted ROWS (
count(*)) and the arithmetic divided that into bytes to get a per-BUCKET figure, then multiplied by a BUCKET count. A rollup holds one row per (server, database, query_hash, sql_handle, bucket), so the units differ by the distinct queries seen per hour — 205x under on a 200-row/hour store, worse on a real fleet box. That is the one direction this preflight exists to prevent, and the code's own comment said so. Fixed withcount(DISTINCT bucket).The suite could not have caught it, which mattered as much as the bug: both live fixtures seeded one row per bucket, making rows and buckets the same number and the error factor exactly 1. That is not a small fixture, it is a fixture of a shape the product never meets. They now seed 12 distinct queries per bucket, and a live test asserts the probe returns a bucket count by ratio rather than by string — so the semantics are pinned, not the spelling.
2. An interrupted backfill left a hole, reported DONE, and the arming gate then armed a purge over it.
min(bucket)was the resume point, the completion verdict, and what the #1680 gate reads before letting the 4-day raw purge drop chunks. Slices ran oldest-first, so the first slice drove that single number to its final value and every later slice was invisible to all three. Proven by execution: killed at slice 32 of 43, 47 of 264 buckets missing, re-run printed "nothing to do — DONE" and exited 0, and the gate then armed over a window whose only surviving copy was the raw rows about to be dropped. The slice-failure path reached the same state, and the "Safe to interrupt" line was false.Fixed by reversing the slice order, not by adding per-slice bookkeeping. Descending makes
floor <= raw_oldestimply completeness by construction: the floor can only reach the bottom once the last, oldest slice has run, in any interleaving. An interrupted run leaves it visibly short — a truthful SHORT instead of a false DONE — and the gate stays closed because it reads the same honest number. Mid-slice kills need no extra state either: a refresh commits per batch newest-first, so the floor lands inside the killed slice and re-planning from it re-covers the remainder. I chose this over per-slice verification because it removes the failure mode rather than detecting it — there is no interleaving left in which a hole can be reported as coverage, so nothing depends on a probe being remembered.Plus the non-blocking sanity clamp. One stray epoch-era row produced a 739,825-slice plan that burned CPU indefinitely. Plans past a 10-year ceiling now REFUSE and name the offending timestamp — that is corruption someone has to go and look at, not history to back fill. A refusal is deliberately not a skip:
[REFUSED]on stderr, DONE blocked, non-zero exit, because burying it among the[OK]lines is exactly how it would be missed.Three findings from the gated live leg
None were visible by reading; all three are product fixes, not test workarounds.
55P03concurrent refresh. The verb runs while the service is UP and the aggregate's own refresh policy lands on top of a slice. Reproduced immediately:EnsureContinuousAggregatesAsyncattaches the policy, the policy runs at once, the next slice failed. Retried, bounded, transient-only — every other SQLSTATE still fails fast.22023"refresh window too small". A day-wide slice's ragged tail is narrower than one bucket for a daily rollup, and the range end is a coverage floor or "now", so it lands mid-bucket most of the time. It aborted the whole daily tier. The planned range now closes on a bucket boundary as well as opening on one.min(bucket)cannot answer a question about one slice. Convergence is judged once, at the end, against raw's oldest row.Testing
Full solution rebuild: 0 warnings, 0 errors. Darling suite: 3723 passed, 0 failed (7 skipped — gated on a live SQL Server and the PG-runtime fixture, neither present here).
Gated live legs ran against a real PostgreSQL 18.4 + TimescaleDB 2.28.1 store stood up from the repo's
pg-runtime.zip, on a store built into the exact broken shape: pre-existing history, rollup createdWITH NO DATA, floor well after raw's.darling.json, not a re-implementation of its loop: dry run changes nothing, the real run'sDONEclaim is re-verified against the store per rollup, and a second run reports nothing to do.[REFUSED]on stderr naming it, no DONE, exit 1.query_stats' held retention policy by itself, throughEnsureRetentionPoliciesAsync— the real seam, not its predicate.Mutation table, each verified RED:
Raw→DailyTierCoverage.Unknown— compiles, reads as deliberate, and is the likely way backViewerDataService.DailySummary.cs:57 (passes TierCoverage.Unknown — routes on NO evidence)ViewerDataService.DailySummary.cs:57 (passes no coverage lookup)count(*)), against the R>1 fixtureINCOMPLETE+ names the exact shortfall; live test redKnown limits of the coverage model, recorded not assumed away
The floor is a depth measure, not a completeness guarantee, and
RetentionTierRouter's docs now say so explicitly. Two shapes slip through a one-dimensional floor: a mid-window materialization hole is invisible to it (a service down longer than the refresh policy's 3-daystart_offsetresumes at now-3d and never backfills the skipped interval, whilemin(bucket)goes on reporting the original deep floor), and a window straddling the floor is served partially with no signal to the caller (returning the tier is still the right choice — raw would return less on a healthy store — but the part below the floor is missing).Both are accepted here because both are strictly better than the age-only routing they replace, which served the entire window as EMPTY in exactly these cases. The docs paragraph exists to block one specific inference: "the coverage gate passed, therefore the result is complete." That belief-shape is how #1759 survived as long as it did — the previous comment asserted the view was "correct to query for any window immediately", it read as reasonable, it was pinned by a test, and nobody re-derived it for months.
Tracked as #1791 with both candidate closures scoped (two-dimensional coverage with gap detection; a partial-coverage notice through the existing
CombineNoticesmachinery). Neither is cut-night material — the first has a real cost question on a 5-minute probe cadence, the second is a compose-contract change.Follow-ups filed
Interaction with the other lanes in flight
Verified rather than assumed: #1772 (delta-seed time bound) and #1768/#1767 (payload dimensions) touch neither the routing ladder nor the rollups' materialization — the CAGGs never carried the payload columns, so the dimension work is invisible here, and the delta seeder writes raw rows the coverage probe only ever reads a
min()from. Branched offorigin/devat50012266, which already contains both. Full-suite green above is over the merged tree.🤖 Generated with Claude Code