Repository navigation
Legacy trio off the refresh grid (#3653) - #4186
Conversation
… off the refresh grid Moves query_stats_hourly, procedure_stats_hourly, query_stats_db_hourly and their three legacy dailies out of HourlyAggregates/DailyAggregates into a new FrozenRollupAggregates list (OffGridAggregates pattern, but with no policy builder at all): the ensure sweep still creates a missing one and now actively detaches any refresh policy an existing store still carries on it, and no converge can re-add one since none of the policy-builder lists name it anymore. RollupBackfill.Targets excludes the frozen six so MaterializationHoleTargets' Single() lookup against Hourly/DailyAggregates cannot throw for a view neither list has anymore. Design points 3, 4, 6 and 7 (raw-purge coverage, successor-hourly retention coverage, the CompressionDeferredUntilFreeze removal, and the frozen-daily compression drain rule) are not yet done — see the PR body for the full handoff. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
… band, drain the frozen dailies' policies Point 6: deletes CompressionDeferredUntilFreeze and its .Where filter, so AggregateCompressionTargets is every member of HourlyAggregates, DailyAggregates and BaselineAggregates with nothing subtracted -- 20 (6 + 7 + 7), not 17. The three interval-honest successor dailies move from no compression policy to hours 11-13 on the band; the seven baseline aggregates shift from hours 11-17 to 14-20. The six hourly members and the four pre-existing daily members keep their hours. Point 7, the drain: EnsureAggregateCompressionAsync's converge already left every non-target aggregate's policy alone (confirmed by reading it -- non-targets are filtered out of the state read entirely, so nothing downstream ever touches them). Adds DrainFrozenDailyCompressionPoliciesAsync, called as a last step in EnsureAggregateCompressionAsync: a probe scoped by name to the three frozen legacy dailies that counts every uncompressed chunk on their materializations (not the age-gated eligible_under_*_rule columns, which answer a different question), a pure predicate (a job exists and no chunk is uncompressed means drain), and a remove_compression_policy call per drained daily, logged. The frozen hourlies get no drain; their chunks age out through retention. Re-pins the TimescaleAggregateCompressionTests pins this touches, by derivation rather than new literals, and adds pure tests for the drain predicate's truth table and the probe SQL's naming. 17 of the 22 red tests listed in the brief remain red -- refresh-grid and phase-slot pins in TimescaleSupportTests/RefreshCeilingProvenancePinTests/ RefreshCeilingStalenessTests unrelated to this lane's points, and coverage/raw-gate/ target-count pins in IntervalHonestHourlyRollupTests/MaterializationHoleRepairTests -- left for a follow-up lane per the PR body's handoff. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…at the successors, not the frozen legacy rollups RawTierCoverage's query_stats/procedure_stats rows and RetentionPolicies' three successor-hourly rows still named the six rollups LC froze off the refresh grid. A gate keyed on a frozen relation either finds it empty once its own retention trims it with nothing refilling it (raw purge, blocks forever) or never actually asks whether the real consumer caught up (successor retention, since the legacy daily's floor is fixed and always looks "covered"). Both are now derived from SupersededHourlyRollups and SupersededDailyRollups through two new lookups (SuccessorOf's daily counterpart, and throwing wrappers for callers that already know the relation is superseded), so the two registries and the two gates cannot drift apart again. Re-pinned the one IntervalHonestHourlyRollupTests test that exercised this by name; left the compression-band, refresh-grid, slot and repair-target pins this touches for their respective owners. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
LC-a2 (points 3 and 4: purge and retention coverage)Commit ed49daf (merged with LC-a3's compression work at 1cc0bb2), pushed to What changed
Red proofDirect inspection of commit 02a1f12 ( Pins fixed vs. left (after merging LC-a3, current HEAD, 165 total)Fixed (mine): 1, Left, by class and reason:
Of the 16 remaining, 12 are refresh-grid/slot (LC-a3's, or already fixed by LC-a3), and 4 are repair-target membership debt from points 1/2 that is not explicitly named in anyone's brief here. Build and tests
🤖 Generated with Claude Code |
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…eze's 12-light, 21-minute window (#4186) The freeze moved the legacy trio off the hourly refresh grid and into FrozenRollupAggregates, with their successors taking the three positions the legacy trio held instead of staying appended behind it. That returns the grid to its pre-Q12 shape (twelve light members, 21-minute heaviest window, 1,050 s watch line) rather than Q12's fifteen/18/900. Ten pins across four files pinned the old Q12 numbers; this re-derives each one against the product's own functions and updates the doc comments to narrate the append-then-freeze round trip. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
LC-a5 (refresh grid re-derived after the freeze)Ran out of budget partway through the 15-pin list plus the coordinator's two added pure pins. Ten are fixed, green, and pushed (commit The freeze returns the hourly grid to its pre-Q12 shape: Fixed and green (10)
Jobs that re-phase once on the first start of this build (from Q12's minute to LC's, all thirteen survivors move because the legacy trio's removal reflows every position): Totals by class: TimescaleSupportTests 7/7 fixed, CompressionPhaseAssignmentTests 1/1, BaselineSupplyTests 1/1, IntervalHonestHourlyRollupTests 1/1 (the other 2 failures in that class, Build: Still red (7) — not reached, with what I worked outRefreshCeilingProvenancePinTests (3):
RefreshCeilingStalenessTests (1): TimescaleContinuousAggregateTests (3): None of my fixes touch a live/rig-gated test; the 16 skipped in my run all need Plain-English checker: not run (past budget; nothing but this comment to check, and it's short). No product code was changed — only test files. No files owned by LC-a4 ( |
…y rollup, repair-list pins re-derived, raw gate covers query_stats_db_interval_hourly DailySummarySql.QueriesCteForCagg/QueriesCteForStitchedCagg looked their relation up in TimescaleSupport.MaterializationHoleTargets, which the freeze (LC) emptied of the six frozen legacy rollups. Every daily-summary read that named a frozen view (legacy-only, RollupCoverage.Unknown, or a stitched splice below the successor floor) threw ArgumentException. Added TimescaleSupport.RollupCoverageProbeTargets: the same per-relation descriptor as MaterializationHoleTargets, but over every RollupViews member including the frozen six, with CREATE text sourced from HourlyAggregates, DailyAggregates or FrozenRollupAggregates. Pointed the three DailySummarySql lookups at it. MaterializationHoleTargets itself is untouched, so the repair walk still never sees a frozen view. Re-pinned MaterializationHoleRepairTests' two lookups that named QueryStatsDailyView directly (now frozen) to its live successor via SuccessorDailyOf, re-derived the stale registered-count literal (26 -> 20) and the daily-after-hourly ordering check (now over the live successor pair, not the frozen legacy pair), and added a pin that MaterializationHoleTargets holds no frozen view. Re-pinned the two MaterializationHoleRepairLiveTests that built their whole scenario on the now-frozen query_stats_hourly to use its live successor instead; the one scenario that compared a live unfiltered rollup against a live filtered one no longer has two live sides after the freeze, so it now proves only the half that is still live: a restart-only hour is neither materialized nor reported as a hole. RawTierCoverage gated the raw query_stats purge on query_stats_interval_hourly only, but query_stats_db_interval_hourly also reads raw collect.query_stats directly. Added it through RequireSuccessorOf(QueryStatsDbHourlyView), fixed the pin that asserted it stayed out, and added a derived pin that walks every RawTierCoverage row and checks its coverage set against every non-frozen RollupViews entry sourced from that raw table directly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
…ollup by its lookup MaterializationHoleScanShapeLiveTests and DailySummaryReadShapeLiveTests both looked a frozen legacy view (query_stats_hourly/_daily) up in MaterializationHoleTargets directly, the same crash pattern as the daily summary's own lookup. Re-pinned the scan-shape oracle to the live successor (SeedAsync already refreshes it over the same windows) and dropped the daily-tier case, whose only live target is now frozen; re-pinned the read-shape oracle to RollupCoverageProbeTargets, mirroring the product's own fix. Compiles clean; not yet run against a live rig in this session. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
Four live TimescaleDB tests against the LC freeze: dropping a frozen legacy hourly's chunks leaves its frozen daily's rows and sums unchanged, no start-path converge ever re-adds a frozen refresh policy while every other aggregate keeps one, the raw purge arms off the successor hourlies' coverage even though the legacy hourlies are empty (and stays held when a coverage relation falls short), and a frozen daily's compression policy drains only once every chunk it will ever hold is compressed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
The legacy six moved to FrozenRollupAggregates, left RollupBackfill.Targets and HourlyRefreshPhaseOrder, and stopped getting a refresh policy. Every test that used one of them as a convenient live example now moves to its interval-honest successor; every test that counts EnsureContinuousAggregatesAsync's "ready" total now adds FrozenRollupAggregates.Length; the compression pin moves from 23 to 20 targets, and the successor dailies now get a real compression policy instead of the removed deferral. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
LC-a4 (the daily summary crash, the lookup audit, the raw gate)Item 1: the daily summary crash
Fix: added Red proof: Coordinator's correction reported 54 CI failures, not 19. Confirmed and fixed on the pure side: Item 2: the same bug elsewhereGrepped production Checked Item 3: MaterializationHoleRepairTests (2 red)Both named a frozen I also found and fixed the same pattern in the live classes the coordinator's correction named, which are not at None of these four live methods ran against a rig this session. All compile, and all are mechanical, Item 4: the raw purge gate
Added Fixed TotalsLocal run, pure classes, Deferred: context budget, did not reach the PG rig this session
Next lane on this branch: run the full live suite on a rig and fix whatever the four unverified-but-fixed methods, |
LC-a6 (live tests re-pinned after the freeze)All 11 assigned tests re-pinned, 0 left red. Verified on a fresh rig (port 55984) before pulling the other lanes' pushes. Confirmed the merge builds clean with 0 warnings afterward. RollupBackfillLiveTests (4 of 4 fixed):
SuccessorDailyLiveTests (2 of 2 fixed):
Ready-count formula (4 tests, all the same root cause): IntervalHonestHourlyRollupLiveTests (1 of 1 fixed):
MeasurementContractCensusTests (1 of 1 fixed, pure):
TimescaleSupportTests (1 of 1 fixed):
TimescaleAggregateCompressionTests (2 of 2 fixed):
Verification: Per the "run every class in every file you touch" rule, I also ran the non-live After that check I pulled No new GitHub issues filed. Every defect found was in-lane and fixed directly. 🤖 Generated with Claude Code |
LC-b (live proofs and the full suite)New file:
Between each revert the source was rebuilt and the specific test re-run alone; after all four, Full suiteRan once on head
23 of the 24 failures are on the coordinator's updated known-red list (54 tests, all owned by LC-a4/LC-a5/LC-a6), by class:
One failure was not on the known list: No PG migration. No product code changed (all four reverts above were local-only and restored before commit; What the coordinator should double-check
🤖 Generated with Claude Code |
The A6 freeze raised HeaviestRefreshWindowMinutes 18→21, RefreshPhaseSlotSeconds 1080→1260, RefreshSlotWarningSeconds 900→1050, and removed the legacy trio (query_stats_hourly, procedure_stats_hourly, query_stats_db_hourly) from HourlyAggregates, reducing LightHourlyRefreshCount 15→12. TimescaleSupport.cs doc-comment prose updated: - Margin sentence: 1080→1260-second slot, 184→364 s margin - Live envelope: 183.9→363.9 s slack, 17.0→28.8%, 3.9→153.9 s below watch - Rejected alternative: 840→1020 s, BELOW→ABOVE the ceiling - Watch line: 900→1050 s, 4→154 s margin below ceiling - Final ordering paragraph: restated for ABOVE case - Scope sentence deleted (covered == total, 12==12) - CompressionPhaseMinutes occupancy: 1080→1260 s, 3→6 minutes on table RefreshCeilingProvenancePinTests.cs: - BELOW→ABOVE regex in rejected-alternative pin - Scope sentence pin and its Verify Require() calls retired (A6 closed gap) - Direct lightCensus[3] check added to Verify (keeps 4th group load-bearing) RefreshCeilingStalenessTests.cs: - 952 s run now InsideSlot again (watch line restored to 1050) - Assertions: >= WatchLine → < WatchLine, margin 52→98, band 4→154 s TimescaleContinuousAggregateTests.cs: - Idempotent test: QueryStatsHourlyView → QueryStatsIntervalHourlyView - Stagger test: 9→6 aggregates, 16→13 phase order, 15→12 light count - Distinct test: 5→3 query_stats consumers, 2→1 procedure_stats consumers - Rotation responsiveness check moved outside loop (3 unbounded views means some rotations preserve sub-group order; the map still responds to at least one rotation, which is what proves it is order-based) Full suite: 13750 total, 0 failed, 693 skipped, 0 errors. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
LC-a5-fixCommit: afa1471 on branch What was fixed (7 PURE test failures, all doc-comment prose and test literal staleness from the A6 freeze): TimescaleSupport.cs doc-comment prose — RefreshSlotWarningSeconds and HeaviestHourlyRefreshObservedCeilingSeconds:
RefreshCeilingProvenancePinTests.cs:
RefreshCeilingStalenessTests.cs:
TimescaleContinuousAggregateTests.cs:
Final suite: 13750 total, 0 failed, 693 skipped, 0 errors — run in the worktree against the committed tree. |
…eshCeilingStaleness, TimescaleContinuousAggregate)
…sumer gate RawTierCoverage now requires BOTH query_stats_interval_hourly AND query_stats_db_interval_hourly to cover raw query_stats before arming the purge gate (LC-a4, #3653). Six of the seven failures were tests that only satisfied one of the two consumers; the seventh was an off-by-one in a scan-shape assertion. Changes per test file: QueryStoreCorrectedRollupLiveTests: Replaced the single legacy query_stats_hourly refresh with both successor hourlies, and changed the DROP sentinel to the successor. Comment tags #3653 LC. RetentionReevaluationTests: After refreshing both successor hourlies (to satisfy the raw-coverage gate), also refresh both successor dailies so the hourly policies' own gate does not go from ARMED to RE-HELD (hourly coverage requires the daily consumer to be non-empty). Updated the ARMED log assertion to the new string.Join coverage string. PayloadDimensionLiveTests: Added delta_worker_time = 1 to the seed row so it qualifies for query_stats_db_interval_hourly (which filters WHERE delta_worker_time IS NOT NULL). Changed the single-view refresh to a foreach over both successor hourlies, with and without force. RollupBackfillLiveTests (RollupCreatedOverExistingHistory): Added a force-refresh of query_stats_db_interval_hourly after the backfill slices. Force is required because the slices exhaust query_stats' shared invalidation log; a plain refresh finds no entries. Floor the window_start to the hourly bucket boundary: refresh_continuous_aggregate only materialises buckets whose START >= window_start, so passing a sub-hour timestamp skips the oldest bucket and leaves coverage one row short. Same floor fix applied to InterruptedBackfill. DailySummaryNotCarriedTests: The repair now targets query_stats_interval_daily (the successor), not the frozen query_stats_daily. Seed the successor daily with a D2 hole so the repair has something to fill. Post-repair read uses the stitch-aware RangeSqlFor(Daily, coverageForRepair, R(0)). MaterializationHoleScanShapeTests: The scan shape produces 7 rows (not 6) because the H20 restart row (delta_elapsed_time IS DISTINCT FROM 0, delta_worker_time IS NULL) is counted by the corrected hourly's WHERE clause but excluded by the db-grain companion's additional filter, giving both views a distinct row count. Full suite result (net10.0-windows, rig 55990): 13750 total, 0 errors, 3 failed (CaptureDownChunkOrderTests, ServerListAndSummaryPlanShapeTests, PgTargetAnomalyTests -- pre-existing, not in modified files), 47 skipped. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
LC-live2All 7 live test failures fixed. Full suite run: 13 750 total, 0 errors, 3 failed (pre-existing: CaptureDownChunkOrderTests, ServerListAndSummaryPlanShapeTests, PgTargetAnomalyTests — not in modified files), 47 skipped. Final SHA: 7d181c9 |
Opus round-1 review at 7d181c9 — GO with one Medium fix required and one High design question for the coordinatorPoints 3 and 4 have landed (ed49daf, 8b632d0). The live tests are green. The PR body's "17 still red" and "no live test" lines are stale. High (design question — coordinator must rule before arming)TS:6029–6030: raw purge held on every field upgrade. The raw Only If Q12 and LC ship in the same release, every field store upgrading from the current release has both purges held from its first start with no way to exit on its own. Options:
Medium (small fix — lane must implement before arming)TS:8743: frozen compression jobs lose the raw converge's exemption.
Fix:
Low
Reviewer answers to your four questions
Other checks (clean)
Status: PR stays in draft until (a) coordinator answers the High design question and (b) the Medium fix lands on this branch. Verdict on the draft: HOLD pending those two. |
|
Round-1 security and data-loss review, PR #4186 at 7d181c9.
|
…llup compression cadence (#3653) Three defects in the LC freeze path: 1. FrozenDailyCompressionDrainStateSql scoped to SupersededDailyRollups (3 dailies) but existing stores also have compression policies on the 3 frozen hourlies; expand to FrozenRollupAggregates (all 6) so the drain cleans up hourly policies too. 2. ConvergeCompressionScheduleAsync lacked a guard for IsFrozenRollupAggregate: the six frozen views are removed from AggregateCompressionTargets, so they fell through the existing IsAggregateCompressionTarget guard and had their compression schedule reset to the 1-hour raw cadence on every start. Added an unconditional skip before the existing guard — frozen views are never ours to tune. 3. LogCompressionActivity's summary debug log counted frozen views in the raw-table bucket, misreporting the raw policy count; subtract frozenPolicies from the total. Test: updated FrozenDailyCompressionDrainStateSql test to assert all 6 views; added IsFrozenRollupAggregate_ReturnsTrueForAllSixFrozenViews_AndFalseForOthers to pin the predicate and prove it is disjoint from AggregateCompressionTargets. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
Medium security finding fix — commit 0cb2188 on fix/3653-a6-freezeDiagnosis confirmed. Three defects in the LC freeze path: Fix 1: FrozenDailyCompressionDrainStateSql scope too narrow (line 7943)The SQL used Fix 2: ConvergeCompressionScheduleAsync missing guard (line 8741)Frozen views are absent from Fix 3: Debug log count off (line 10828)
Tests
Build: 0 errors, 0 warnings. |
…loors (#3653 LC) After the LC freeze, callers that use coverage.For(legacyHourly, legacyDaily) to determine whether the hourly tier can serve a window were broken in two shapes: - Fresh store (just upgraded): legacy starts WITH NO DATA -> null floor -> every window older than ~4 days falls to raw. - 90+ days post-freeze: retention drains the legacy -> same null-floor collapse. Fix: RollupCoverage.For() now stitches the legacy and its interval-honest successor (from TimescaleSupport.SupersededHourlyRollups / SupersededDailyRollups), taking the deeper (earlier) of the two non-null floors as the effective floor for each tier. The stitch is backward-compatible: when only the legacy has data, the legacy floor wins unchanged. All seven production callers (FinOps.Workload.cs:247/408/476, ViewerDataService.DailySummary.cs:83, QueryTrends.cs:508, DarlingHealthReader.cs:420, DarlingMcpTrendTools.cs:383) get the correct behavior without individual edits. HourlyRelationFor comment updated: the legacy is no longer guaranteed deeper "by construction" on post-freeze stores; the stitched floor from For() is what makes the tier decision correct. Adds four pure regression tests covering the fresh-store, post-trim, pre-freeze, and stitch-boundary shapes. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
High fix: raw purge gate holds forever on field upgrade (#3653)Implemented Option 4 (stitched coverage) plus the zero-interval source filter. Pushed to Problem confirmedTwo independent issues combined to hold the purge gate open permanently on a field-upgrading store:
What changed
(SELECT COALESCE(LEAST(l.mn, s.mn), l.mn, s.mn)
FROM (SELECT min(bucket) AS mn FROM collect.{legacy}) l
CROSS JOIN (SELECT min(bucket) AS mn FROM collect.{successor}) s)When the successor is empty: The stitching is driven by a new Source filter for Tests
Not coveredThe live test runs against a PG rig and is marked skip when |
3cfeeea's stitched coverage (frozen legacy + successor) counted legacy+successor as covering raw with no gap check, and its doc claimed the stitch was gap-free by construction. That is false: a store stopped 1-4 days before the successor's first refresh leaves a raw tail (about 2h) between the legacy's last bucket and the successor's floor that neither side ever materializes; RawRetentionInterval (4 days) then purges it permanently. RetentionArmSafetySql now probes raw itself for a filter-admitted row in that seam (legacy's last bucket to the successor's floor, or +infinity when the successor is empty) before trusting the stitched floor. A row found falls back to the successor's own min(bucket) (Short, honestly) instead of the false Covered; an empty seam keeps the old stitched LEAST(legacy.min, successor.min), simplified from the redundant COALESCE(LEAST(...), ...) since PostgreSQL's LEAST already ignores NULLs. A floors-only gap test was rejected: it would deadlock forever on an outage with no rows in the seam or a tail already purged. RepairMaterializationHolesAsync's hole walk now scans a successor with a frozen legacy from the legacy's last bucket rather than the successor's own floor, so the seam tail is repaired on the next start and the gate's fallback releases on its own. Also fixes FieldUpgrade_EmptySuccessors_LegacyFilled_ReportsRawPurgeCovered, which was already failing on this branch's own CI (confirmed before touching it): its fixture refreshed only two of query_stats' three legacies, leaving the LC-a4 db-grain pair fully empty on both sides regardless of this change. Adds the outage-shape and empty-seam live tests plus a pure SQL-shape pin test; corrects the "gap-free by construction" doc claims in both files. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
Seam-gap fix for the stitched raw purge gate (commit 3098a9e on this branch)Pushed directly to The defect this closes3cfeeea's stitched coverage (frozen legacy + successor) counted Fix(a) Gate ( (b) Hole walk ( (d) Docs: replaced the false "gap-free by construction" text in both files' doc comments with the seam Tests (all run locally; see "Rig used" below)
Targeted classes ( The 2 failures are Rig disclosure (please read)I built a private rig before the coordinator's port/dir assignment arrived (copied an existing extracted TrailersThis push's commit carries the session trailer from this conversation's opening system reminder CHANGELOGNone added. This PR is still an undeployed, in-progress part of #3653 (draft, base Files changed
🤖 Generated with Claude Code |
RepairMaterializationHolesAsync's seam bound (3098a9e) still clamped the scan to `from = max(seamFloor, horizon)`. For a store stopped more than the raw span (4 days for query_stats/procedure_stats) before its first start on this version, the seam tail lies below that horizon and was never scanned, so RetentionArmSafetySql's seam probe keeps finding the un-repaired rows and the raw purge gate holds forever without a manual --backfill-rollups. MaterializationHoleScanWindows is a new pure helper: the ordinary window stays exactly [max(floor, horizon), ceiling], and a seam window [seamFloor, floor) is added whenever a seam exists, scanned in full however far below the horizon it reaches. RepairMaterializationHolesAsync now scans each window and concatenates the holes before merging and capping, unchanged from there. Adds pure tests for the new helper (no seam, seam above the horizon, seam below it, the ordinary window never starting below the horizon) and a 6-day-outage live twin of Outage_SeamBetweenFrozenLegacyAndSuccessor_ HoleWalkRepairsItAndGateReleases, which only ever exercised the 2-day (short) case. Corrects both doc comments' "releases on its own" claim to state the automatic release is bounded by the per-start repair cap, and the AggregatesSkipped doc's horizon-skip reason to exempt a seam window. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
FrozenRollupAggregates' doc said "the four still with a drop_chunks policy"; RetentionPolicies names three (the legacy hourlies), not four. The aggregate setup loop's catch logged "composer queries fall back to raw scans" for every failure, including a frozen rollup's own refresh- policy detach. That view's CREATE already succeeded earlier in the same try, so it keeps serving composer queries same as ever on a detach failure — it just stays on its pre-freeze refresh schedule instead of frozen until a later start retries the detach. Branch on IsFrozenRollupAggregate and log that case truthfully; every other view keeps the existing message. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
Seam scan follow-up: the seam now repairs even when the outage outlasts the horizonPushed to The defect
The fixNew pure helper
Corrected both doc comments that claimed unconditional automatic release: the class doc's seam paragraph and Coordinator's two review items (separate commit,
|
The #4186 Low fix reworded the frozen-view branch of the continuous-aggregate setup catch to talk about the refresh-policy detach only. The same catch also sees a failed CREATE (a fresh store) or the width step, so the line now names both steps instead of claiming the view exists. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ
Round-3 review at 35a0aefScope: the 7 commits after 7d181c9 ( Terms used below:
Paths are under Tally: 2 High, 2 Medium, 3 Low. Q2 found a gate that can release early (H1). Q3, Q4 and Q5 found nothing. H1. A partial seam repair moves the successor's floor below unrepaired seam rows, and then the gate reports CoveredWhere:
The seam lies below the successor's floor. So the oldest seam bucket that a repair writes becomes the new
Sequence:
A narrow seam can also cause the same release if it has two or more ranges. Suppose the older range is repaired and then a newer range fails. The refresh can throw, and then the catch at line 624 skips the rest of the target. A shutdown can cancel it. It can also leave buckets, and then Trigger: in the three shapes the brief names, the seam is only the tail of about 2 hours. That tail is one range of at most 3 buckets, so the cap never splits it. The cap trigger needs a refresh that stalls for about 2 days or more while collection continues. The stall can be the legacy's refresh before the upgrade. It can also be the successor's first refresh after the upgrade, for example when its policy failed to attach at the upgrade start. The failure trigger needs a seam with two or more ranges. Fix: repair the seam window newest-first. Take its cap from the top of the seam (the buckets next to Then the floor only grows down without a gap, and every unrepaired seam row stays inside the probe window. This is the property that H2. The stitch trusts the frozen legacy's own span, including holes from before the upgrade that nothing will repairWhere:
Sequence:
Before 3.8.0 loses the same rows on the same schedule, so this is not a regression against the released build. It meets the brief's High rule because the gate reports Covered over rows that no rollup holds. Fix (a design decision): use one of these two options.
M1. An armed raw purge runs when PostgreSQL starts, and it drops a seam older than 4 days before the service can hold or repair it (pre-existing)Where:
Sequence: a 3.8.0 store with an armed raw purge has PostgreSQL stopped for 6 days. When PostgreSQL starts, the TimescaleDB scheduler runs the overdue retention job for
These commits do not cause this, and 3.8.0 loses the same tail. So I rank it Medium, although it meets the letter of the High rule. The problem for this PR is a claim. The doc for Fix: hold the raw retention jobs in the service's graceful shutdown path, and let the start sweep arm them again after it measures. This does not cover a crash or a stop of PostgreSQL alone. A full fix makes each drop run check coverage first, as a custom job. We recommend a separate issue for that, and a sentence in the doc that states the limit. M2. After an outage of more than a day across the upgrade, the gate holds until the second startWhere:
Sequence: a 3.8.0 store is stopped for 2 days and then takes this build. At the first start, the successors are new and empty, so the walk skips them, and the gate reads Short. After the successor's first refresh, the seam holds the tail of about 2 hours, and each hourly evaluation reads Short. Nothing repairs the tail until the next start. Until then, the raw purge for both tables stays held, and raw grows by a day of collection each day. No rows are lost. But a store that seldom restarts needs a manual step (a restart or Fix: run only the seam window, with the same cap, from the hourly retention evaluation when the probe finds a seam. Another option is to run the walk once more after the successor's first materialization. If you choose neither, change the doc to say that the release needs a second start. L1. Every upgrade from 3.8.0 holds the raw purge again, with a warning that tells the operator to backfillWhere: Sequence: a healthy 3.8.0 store has an armed purge and no outage. At the upgrade start the successors are empty, so the probe window runs from The successor's first refresh clears it within about 2 hours. No rows are lost. But the warning sends the operator to a heavy verb for a state that clears by itself. Fix: when L2. A repaired seam older than 3 days never reaches the successor daily, so the successor hourly's 90-day retention is heldWhere:
Sequence: after a 5-day outage, the second start repairs the seam into No rows are lost, because the hourly keeps them. Fix: give each successor daily a seam window that starts at its hourly's floor. Or make the seam repair also refresh the dependent daily over the repaired days. L3. The probe filter for the db-grain slot is narrower than that successor's WHEREWhere: A seam row with a NULL delta and a non-zero interval makes the probe find a row that the walk never treats as a hole. The purge then holds forever. This fails in the safe direction. The current collector writes non-null deltas ( Fix: build the filter for each slot from Answers by questionQ1: yes, see H1, H2 and M1. The zero-interval filter from A fresh install is safe. The frozen views stay empty, so Q2: yes, see H1. Q3: nothing new. A seam bucket does not exist before the repair, so compression cannot reach it first. The seam refresh writes below the successor's floor, where no compressed batch exists. The code has its own measurement on 2.28.1 ( A refresh that stops short does not fail quietly. The walk scans again, forces a refresh, and logs what remains ( One point is unverified. The walk's refreshes do not lift Q4: nothing. The state probe reads only the six frozen names in A frozen view keeps a running policy until every chunk is compressed, by design. Nothing refreshes a frozen view, so it gets no new chunk. A paused job never compresses, so the drain never removes it. A missing job leaves nothing to drain. Both of these cost disk only. Compression and policy removal remove no rows. Q5: nothing for purge or retention. Every caller of
The purge and retention paths use only
|
|
H1: a fix lane is running with the ruled fix (the seam walks newest-first and stops at the first failed range) and a live test. H2: ruled parity with 3.8.0. The same lane states that in the |
…4186 round-3 H1) A seam wider than one start's repair cap, or with two or more ranges, could move the successor's floor past a still-unrepaired seam bucket: oldest-first capping and repair let the OLDEST buckets close first, which drags the floor (a bare min(bucket)) down to them even while newer seam buckets in between stay holes, and RetentionArmSafetySql's probe stops looking above the new floor. The raw purge could then arm and drop rows neither rollup ever held. Fix: the seam window's holes are now scanned, capped and walked separately from the ordinary window's. The seam takes its cap from the newest end (closest to the successor's floor) and walks descending; the ordinary window is unchanged (oldest-first, same cap, same continue-past-a-remainder handling, since an interior repair can never move the floor). The seam walk stops at the first range that throws (propagates to the existing per-target catch) or leaves buckets standing (remaining > 0) rather than moving on to an older range. Both windows share one cap budget per aggregate per start, seam first, so "a start never re-materializes more for one aggregate than an ordinary policy run does" stays true. CapMaterializationHoleRepairs gains a newestFirst parameter (default false, unchanged behavior) that also splits a straddling range at its older edge instead of its newer one, so the kept portion stays adjacent to whatever is already materialized. Also corrects RetentionArmSafetySql's doc: it claimed an over-horizon seam "still gets repaired" and the gate "releases automatically, with no manual step" unconditionally. Both are true only when the raw purge was already held when the store stopped (#4299) and the outage did not cross the upgrade (#4300) respectively. Adds the H2-parity note: the stitch trusts the frozen legacy's own span without checking it, which can read Covered over a pre-upgrade hole nothing here repairs -- accepted because 3.8.0 loses the same rows on the same schedule, and the A6 backfill (--backfill-rollups) covers it for an operator who runs it. Two live tests against a real TimescaleDB rig: - a 35-bucket seam (over the 24-bucket cap) stays Short after one walk (top 24 repaired, bottom 11 still a hole) and reads Covered only after a second walk closes the rest; revert-proved by temporarily restoring oldest-first, which fails at the same assertion. - a two-range seam where the newer range's refresh is forced to raise 23514 via a CHECK constraint on its own materialization chunk: the older range is never touched and the gate still reads not-safe. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
Round-3 H1 fix, pushed to
|
…floor The walk's #4186 seam-fix block lowered seamFloor to the frozen legacy's last bucket + one bucket width. Per the ruling (#4301 comment 5844202704), the walk now lowers it to raw's own filtered floor (AlignDown'd), the same bound RetentionArmSafetySql's gate probes from, whenever that reaches further back than the successor's own floor. Everything downstream (the successor-only scan, the newest-first cap and walk, the shared cap with the ordinary window) is unchanged and is itself the contiguous downward fill: RollupCoverage.StitchedRelationSql splits its read at the successor's floor, so every row below it must land as a successor bucket for the stitch to read each row exactly once. Removed the now-dead legacyInteriorFrom/legacyInteriorTo third-window parameters from MaterializationHoleScanWindows and the two pins that exercised them (ScanWindows_LegacyInterior_*, ScanWindows_NoLegacyInterior_*): under this ruling a legacy-interior hole is repaired by the same newest-first seam descent as everything else below the successor's floor, so the separate oldest-first interior branch never gets built. Removed LegacySuccessorHoleScanSql (the walk's list-form twin of the gate's LegacySuccessorHoleExistsSql) since the walk no longer shares that hole definition; RetentionArmSafetySql keeps LegacySuccessorHoleExistsSql for its own probe.
…the successor down to raw's floor (#4401) Closes a gap in the raw-purge gate after the legacy rollups were frozen (#4186). - RetentionArmSafetySql no longer trusts the frozen legacy rollup's whole span. It checks for a hole inside that span, not only for a seam above it, before reporting the raw purge covered. - The repair walk fills the successor rollup contiguously downward to raw's own floor, newest first, instead of stopping at the legacy's last materialized bucket. An isolated repair below the legacy's last bucket would move the successor's first bucket down, so stitched reads would take the hours in between from the successor, which doesn't hold them. - FrozenRollupLiveTests has five new pins. Three existing seam tests have their expected counts and floors updated for the contiguous fill; no assertion was loosened. Refs #4301
This PR freezes six legacy rollup views so they stop refreshing:
query_stats_hourly,procedure_stats_hourly,query_stats_db_hourly, and their three hierarchical dailies. Their interval-honest successors (Q12 hourlies, A6 LB dailies) are already shipped and already cover the same history. The freeze moves purge and retention coverage onto those successors. It also fixes a read-routing regression the freeze itself caused. It closes a gap in the raw-purge gate too, at the seam between the frozen legacy data and the successor's own history.Built in several passes on
fix/3653-a6-freeze. Two review passes, a design review and a security and data-loss review, ran at commit7d181c99. Every High and Medium finding they raised is fixed on the branch. A third pass then reviewed those fixes at commit35a0aefb. Its High, a way for the purge gate to release early, is fixed in commitc3291f66. Its other findings are written up in the code as accepted limits or tracked in follow-up issues (see Review).Part of #3653
What changed
The freeze (commit
02a1f124). Six legacy rollups move out ofHourlyAggregatesandDailyAggregates. They go into a new list,FrozenRollupAggregates:query_stats_hourly,procedure_stats_hourly,query_stats_db_hourly, and their three hierarchical dailies.EnsureContinuousAggregatesAsyncstill creates all six, on a fresh store or one upgrading past this PR. But it detaches any refresh policy an existing store already has on them, every start, and it never attaches a new one.RollupBackfill.Targets, and the materialization-hole repair targets built from it, exclude the six. Nothing ever backfills or repairs a rollup that no longer advances. Emptying that list broke a separate lookup inDailySummarySql, which then threw for any read naming a frozen view. Commit8b632d0badded its own lookup,RollupCoverageProbeTargets, which still includes all six, and fixed it.Purge and retention coverage moved to the successors (commit
ed49daf6).RawTierCoverage'squery_statsandprocedure_statsrows named the frozen legacy hourlies as their raw-purge coverage. They now name the interval-honest successors instead, through a newRequireSuccessorOflookup.query_stats's row also namesquery_stats_db_hourly's own successor, added in commit8b632d0b.query_stats_db_interval_hourlyreads rawquery_statsdirectly too, so it needed the same coverage.RetentionPolicies' three successor-hourly rows named the legacy daily as their own retention coverage. They now name their own successor daily instead, throughSuccessorDailyOf.The stitched floor: purge gate and read routing (commits
3cfeeea2,0174f585). A successor's own history starts empty on an upgrading store. That broke two things independently. The raw-purge gate held forever, with no self-exit. Every FinOps, health and MCP reader that resolves a tier throughRollupCoverage.For()fell back to raw, for any window older than about 4 days. Both are fixed the same way. Each stitches the legacy floor onto the successor's own floor, and takes whichever floor is deeper (earlier). The gate does this inRetentionArmSafetySql. Reads do it inRollupCoverage.For().The seam probe and the hole-walk repair (commit
3098a9e5). The stitch above was not gap-free. An earlier version of its own doc comment claimed it was, and that claim was false. A store can stay stopped for more than a day across the upgrade. It then leaves a raw tail between the legacy's last bucket and the successor's floor. The tail is as long as the stop, minus the one day that the successor's first refresh reaches back. Neither side materializes that tail, and without a repair the raw purge deletes it.RetentionArmSafetySqlnow probes raw itself, for a row in the tail, before it trusts the stitched floor.RepairMaterializationHolesAsync's scan now reaches back to the legacy's last bucket too, not just the successor's own floor. The tail then reads as an ordinary hole. The service repairs it at startup, up to 24 hourly buckets per start. After an outage of more than a day across the upgrade, the repair starts at the second start (#4186 gate follow-ups: release needs a second start after an outage, plus three smaller holds #4300). Commitc3291f66makes the repair walk the tail from its newest bucket down, and stop at the first range that fails. Before that, the walk took the oldest buckets first. A partial repair then moved the successor's floor below tail rows that were still missing, and the gate read Covered too early. Commitc14b46a0gives the tail its own scan window. Before it, the scan stopped at a 4-day horizon, so a tail older than that was never repaired, and the gate held for good.The zero-interval source filter (commit
3cfeeea2). Raw's very first collection row, on any store, always hassample_interval_seconds = 0. The successors already exclude that row from their own history. The gate's ownsource_oldestcheck now excludes it too. That one row can no longer hold the gate open past where the successors themselves start.The compression band (commits
4beeb261,0cb21885).CompressionDeferredUntilFreezeis gone. That was the set that held the 3 successor dailies out of the compression band, until this freeze shipped.AggregateCompressionTargetsis 20 members now: 6 hourly, 7 daily, 7 baseline, nothing excluded. (devhas 23 today. The legacy trio still counts there, and the 3 successor dailies are still deferred.) The legacy trio also leftHourlyAggregatesandDailyAggregates, not just the compression list. So every one of the 17 compression policies that exist today moves to a new band hour on the first start after this ships. Each is counted fromAggregateCompressionBandFirstHour. No aggregate'scompress_afterwindow changes. Only its hour does. The 3 successor dailies,query_stats_interval_daily,procedure_stats_interval_dailyandquery_stats_db_interval_daily, are newly compression-registered, at hours 11 to 13.FrozenDailyCompressionDrainStateSqlnow scopes to all 6 frozen views, not just the 3 dailies.DrainFrozenDailyCompressionPoliciesAsyncthen drains a pre-freeze store's frozen HOURLY compression jobs too, once every chunk on that view is compressed.The Medium review fix (commit
0cb21885).ConvergeCompressionScheduleAsyncruns every start. It only skips views still listed inAggregateCompressionTargets. The frozen six left that list. So an existing store's once-a-day compression jobs on them were being reset to the raw tier's 1-hour cadence, on every start. The converge now skipsIsFrozenRollupAggregateviews first, unconditionally, before that check. The same commit also widened the drain's view filter to all 6 frozen views (above). It fixed a debug log too, which was still counting the frozen views inside the raw-table bucket.Review
Two review passes ran at commit
7d181c99: a design review and a security and data-loss review. Every fix below came after both passes. Each fix was checked against the code; no separate reviewer has read them.High:
query_statsandprocedure_statsheld on every field upgrade (design review). A successor's history starts 1 day before the successor exists, raw keeps 4 days, and only--backfill-rollupsreleased the gate. The fix is a stitched floor with no gap. Commit3cfeeea2stitches the legacy floor onto the successor's floor inRetentionArmSafetySql. Commit3098a9e5adds the seam probe and the hole-walk repair. Commitc14b46a0fixes a gap in3098a9e5found during the fix check: its repair scan was still limited to the 4-day horizon.sample_interval_seconds = 0, and the collector writes 0 on its first pass (security review). Commit3cfeeea2applies the same filter to the gate'ssource_oldestcheck.RollupCoverage.For()read the legacy floor alone (security review). Windows older than about 4 days fell back to raw. Commit0174f585stitches the floors there the same way the gate does.Medium:
ConvergeCompressionScheduleAsyncruns every start and skips only views inAggregateCompressionTargets. The frozen six left that list, so their once-a-day compression jobs were reset to the raw tier's 1-hour cadence on every start (design review). Commit0cb21885skips frozen views first. The same commit widens the drain to all 6 frozen views and fixes a debug log that counted them as raw tables.Low:
drop_chunkspolicy. It is three. Fixed in commit22cbb028.22cbb028. Commit35a0aefbthen made the warning name a failed create too, because the same catch sees one.0cb21885's wider drain covers them now.decompress_chunkrun between the drain's check and its removal can leave one chunk uncompressed. That costs storage, not data. No change needed.Round 3, at
35a0aefbA third pass read the fix commits and found 2 High, 2 Medium and 3 Low findings.
c3291f66fixes it (above). The fix was checked against the code, and no fourth review round runs.RetentionArmSafetySqldoc says so, and names--backfill-rollupsas the way to fill the successor down to raw's oldest row. Raw purge gate: check the frozen legacy's span bucket by bucket, and repair holes inside it #4301 tracks a full fix.Known limits
main, 2026-09-17) has no hole walk. A store that upgrades straight from it keeps any hole inside the legacy rollups from an outage in its last 4 days. After the freeze nothing repairs a legacy rollup, and the raw purge then deletes those rows. A store on adevbuild with the hole walk repaired such holes at each start, up to 24 buckets per view per start.query_statsandprocedure_statsstays held, and raw keeps growing.--backfill-rollupscloses it in one run.Tests
Four live tests in
FrozenRollupLiveTests.csdrive the product's own start-up entry points against a real TimescaleDB instance, not a re-implementation of them:DroppingFrozenHourlyChunks_LeavesItsFrozenDailyUnchangeddrivesEnsureContinuousAggregatesAsyncandRepairMaterializationHolesAsync.StartPathRunTwice_NeverReAddsAFrozenRefreshPolicy_AndEveryOtherAggregateKeepsOnedrivesEnsureContinuousAggregatesAsync, run twice.RawPurge_ArmsOffSuccessorHourlyCoverage_NotTheEmptyLegacyOnesdrivesEnsureRetentionPoliciesAsync.FrozenDailyCompressionPolicy_DrainsOnlyOnceEveryChunkIsCompresseddrivesEnsureAggregateCompressionAsync.Each has a revert-proof. The author reverted the guarding code, watched the test fail with the expected error, then restored the code and watched it pass. The seam fix (commit
3098a9e5) adds two more live tests the same way,Outage_SeamBetweenFrozenLegacyAndSuccessor_HoleWalkRepairsItAndGateReleasesandOutage_EmptySeam_GateReportsCoveredWithNoHoleWalk, plus a pure pin,RetentionArmSafetySql_StitchedSlot_ProbesSeamBeforeFallingBackToLegacyFloor. All three carry the same revert-then-restore proof. A fourth live test from the High fix,FieldUpgrade_EmptySuccessors_LegacyFilled_ReportsRawPurgeCovered, needsDARLING_TEST_PGto run and has no revert-proof recorded in the comments.Commit
c3291f66adds a pure pin,TheCap_NewestFirst_TakesTheNewestFirst_SplitsAStraddlingRangeAtItsOlderEdge_AndDefersTheRest, and two live tests:Outage_SeamWiderThanTheCap_NewestFirstRepairsTheTopAndKeepsTheGateHeldUntilFullyRepaired: a 35-bucket seam keeps the gate held after one walk, and releases it after the second. With the old oldest-first order, it failed at the first gate check.Outage_SeamTwoRanges_NewerRangeFails_OlderRangeUntouchedAndGateStillNotSafe: aCHECKconstraint makes the newer range's refresh fail. The older range is never touched, and the gate stays held.The changed classes were run, all green, but not the full suite. CI runs the full suite.
Commit
c14b46a0adds five pure tests for its newMaterializationHoleScanWindowshelper, and a live twin of the seam test for a 6-day outage,Outage_SeamOlderThanTheHorizon_HoleWalkStillRepairsItAndGateReleases. With the old horizon limit put back, the new live test fails ("expected the 2-bucket seam tail to be repaired, got 0"). With the fix restored, it passes. On a fresh local rig,FrozenRollupLiveTests(8 tests),TimescaleContinuousAggregateTests(41) andMaterializationHoleRepairTests(16) passed with no failures, before and after a merge ofdev. The full suite did not run locally after this commit. GitHub Actions CI passed on35a0aefband again on the merged headfd53564c: build, Darling PostgreSQL tests and Lite tests.Every test the freeze broke is fixed. This body drops the original test-status inventory, because none of it is current any more.
The last full local run of
Darling.Testswas on3098a9e5, before the horizon fix. It reported 13760 total tests: 0 errors, 2 failed, 47 skipped, 1 not run, in 553 seconds. Both failures,ServerListAndSummaryPlanShapeTestsandCaptureDownChunkOrderTests, are chunk-visitation-count plan-shape checks in a subsystem this PR does not touch.Whether both are pre-existing was not confirmed locally. The CI run against the merged dev (see the merge gate) is the check.
Merge gate
The earlier prerequisites are merged: LA (#4182), LB (#4181) and LA-8 (#4184).
docs/runbooks/a6-successor-daily-backfill.md).min(bucket)ofquery_stats_interval_hourlyand ofquery_stats_db_interval_hourlyis at or before rawquery_stats' oldest row.7d181c99, and every High and Medium they found is fixed. A third pass ran at35a0aefb. Its H1 is fixed inc3291f66, and that fix was checked against the code. Its H2 is accepted as 3.8.0 parity, and #4299, #4300 and #4301 track the rest. No fourth round runs.CHANGELOG entry
SECTION: Changed
ENTRY:
query_stats_hourly,procedure_stats_hourly,query_stats_db_hourlyand their daily rollups keep the history they already hold. Their interval-honest successors take over. Reads older than the successors' own history still use them. New data goes only to the successors, so the store no longer spends refresh time on the old rollups. The rawquery_statsandprocedure_statspurge now waits until the successors and the old rollups together cover every raw row. After a long outage, the service fills the gap at startup, up to a day of it per start. Only then does the purge resume.--backfill-rollupsfills it in one run. Existing compression jobs move to new hours once. Dashboards, alerts and queries do not change.REF:
[Legacy trio off the refresh grid (#3653) #4186]: Legacy trio off the refresh grid (#3653) #4186
This entry goes into
CHANGELOG.mdin the batched splice after this PR merges.