Skip to content

Three hourly rollups counted every restart's zero as a sample and a >24 h outage left a two-hour hole under a floor that said covered: interval-honest successors beside the legacy trio with the phase grid re-derived by its own method, and a targeted refresh of each materialization hole at service start (partial #3653 — items 16 + 9, Q12 + Q10) - #3731

Merged
erikdarlingdata merged 2 commits into
devfrom
fix/3653-hourly-successors-hole
Sep 19, 2026

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 19, 2026 •

Copy link
Copy Markdown
Owner

Partial #3653 — items 16 + 9 (Q12, Q10)

Two commits, one lane, the same files. The first retires the last three members of the measurement contract's rule-5 roster the way Erik ruled (Q12: three interval-honest hourly successors on the #3698 pattern, the phase grid re-derived by its own method). The second settles item 9's last clause the way Erik ruled (Q10: a targeted refresh of the hole at service start, not a wider start_offset).

Commit 1 — the lie, and the three successors (item 16, Q12)

Mechanism. A restart writes a (0, 0) row: every delta 0 beside sample_interval_seconds = 0, the collector's own "not knowable" verdict. query_stats_hourly, procedure_stats_hourly and query_stats_db_hourly aggregate it — count(*) AS sample_count counts it as a sample on all three, min(delta_*) reads it as a real floor on the two that carry a min. A continuous aggregate's query cannot be altered, and a DROP + recreate forfeits the 90-day tier and cascades the indefinite daily tier hierarchical from it.

What lands. query_stats_interval_hourly, procedure_stats_interval_hourly, query_stats_db_interval_hourly: identical dimensions and column names, the verdict baked into the WHERE (sample_interval_seconds IS DISTINCT FROM 0; the db-grain successor keeps its legacy's delta_worker_time IS NOT NULL beside it), the measured interval carried as sum(sample_interval_seconds) AS sample_interval_seconds_sum (a per-bucket SUM rather than the baseline pair's max, because this rollup groups an hour of collections and a hierarchical daily can sum it again), WITH NO DATA, materialized_only like the legacy for the #1759 coverage-probe reason. Registered in HourlyAggregates, RollupViews (so they are on the --backfill-rollups plan, in the coverage probe and under the supply rule), RollupProbeSql/RollupAvailability (+3 flags, WithoutIntervalHourlies for the pre-this-build store shape), and RetentionPolicies at the hourly tier's 90-day horizon under the leaf rule (coverage = the sibling legacy daily over the same source — the corrected Query Store hourly's shape exactly).

Why this shape and not #3698's replace-in-place — the load-bearing departure. Each legacy has an indefinite daily tier hierarchical FROM it (query_stats_daily etc.), and a continuous aggregate's source is fixed at CREATE. Dropping the legacy cascades the daily; removing its refresh policy freezes the daily at its watermark and serves every daily-tier window ending at now incomplete — the #1759 shape. So the legacy trio stays registered, refreshing, phased, retained and compressed for good under the current reader model, and the successors are appended. SupersededHourlyRollups names the three (legacy, successor, dependent daily) and its essay states what that costs and what the daily tier does NOT get: the dailies inherit the legacy's contamination at the day grain until they have successors of their own, and that lane must re-derive the daily compression band first — it is now FULL at 23 of 23 hours (AggregateCompressionBandHourFor throws on a 24th member). The raw purge is deliberately NOT gated on the successors (the #1661 query_stats_db_hourly precedent, not the #1849 L1 one): gating would hold every store's query_stats/procedure_stats purge from the first start until an operator backfilled.

The grid, re-derived by the documented method — the numbers. Appending three members: LightHourlyRefreshCount 12→15, LightBandSpanMinutes 11→14, HeaviestRefreshStartMinute 15→18, HeaviestRefreshWindowMinutes 21→18, RefreshPhaseSlotSeconds 1,260→1,080, RefreshSlotWarningSeconds 1,050→900. The ceiling HeaviestHourlyRefreshObservedCeilingSeconds = 896 sits 4 s under the line — the ceiling essay's own "18 whole minutes is the smallest window whose five-sixths line clears 896 s", reached exactly. The pinned ordering ceiling < warning < slot holds. The compression band's minute (:35) and minutes (:36–:59) do not move (the band is the hour's remainder; the window shrank by what the light band grew). Positions: the unbounded-cardinality successor takes 12 (the fourth every-fourth slot); the two bounded successors are dealt into the bounded class ahead of the baselines (3 and 5), so the seven baseline policies move two positions later each — nine sub-3 s policies re-phased once by the converge. Every changed literal in TimescaleSupportTests, RefreshCeilingProvenancePinTests, RefreshCeilingStalenessTests, TimescaleContinuousAggregateTests, TimescaleAggregateCompressionTests, BaselineSupplyTests carries the derivation beside it; the provenance pins parse the constant's prose and every stated figure moved with them.

Three consequences of that grid are pinned rather than found:

  1. The 900 s line coincides exactly with Query Store aggregate tax scales serially with fleet size — measure, alert, and evaluate before large onboardings #2136's cadence knob (25% of 3,600 s; both compare >=), so a run at or past it raises both signals on the same reading — the correct remedy no longer arrives second.
  2. The 952 s clean run of 2026-09-08 17:00Z that RefreshCeilingStalenessTests quotes as the falsification of the 896 s constant now classifies ApproachingSlot (was routine). A recurrence of that shape logs a warning, 128 s inside the 1,080 s wall. That case was written to expire; it expired from the other side and was re-read, not re-typed.
  3. The rejected alternative watch line (slot less one guard band) is back below the ceiling (840 s vs 896), so RefreshSlotWarningSeconds rejects a 480s alternative on 'three of the five' readings — only one of the five listed readings reaches 480s, and that count is the sole stated reason #3107's ordering rejection is true again beside Re-derive the refresh grid so contention does not depend on list position: no step works, because Other*4 < Heavy has no step term #3174's coupling one; the paragraph and pin re-taken.

Readers. One supply rule — TimescaleSupport.PrefersSuccessor, the rule PgBaselineProvider already applied per server, moved to Storage and aliased there (not duplicated) — applied to a store's materialized floors by RollupCoverage.HourlyRelationFor(legacy, windowStart) AFTER the tier is decided over the legacy pair (the deeper of the two on any upgraded store). An absent successor is never named (the coverage now carries the availability it was probed with; Unknown and the two-argument constructor answer the legacy). Wired through: the compose router (ResolveFamily), the MCP duration-trend route (DurationTrendRoute.HourlyView is the resolved relation; the hourly SQL is built from it), get_query_trend's keyed history (QueryHistoryHourlySqlFor), the daily summary (DailySummarySql.RangeSqlFor(tier, relation) — calendar and get_daily_health make the identical call), the three FinOps workload readers. LogSupersededHourlyRollupCoverageAsync reports the hand-over on every start (an instrument, not a mechanism).

Census. Rule 5's roster keeps its five names — the legacy CREATE text still exists (registered here; the live fixture there) — and gains the stronger invariant: every member is superseded by a registered successor carrying the predicate, pair by pair. The brief's "roster empties" could not be true and is reported as such.

Commit 2 — the >24 h hole, repaired at start (item 9, Q10)

Mechanism (the #3698 honest read). A refresh policy re-materializes [now - start_offset, now - end_offset]. Down longer than the hourly window, the first refresh after resume opens past the raw collected between the last pre-outage refresh's window end and the moment collection stopped — about two hours — and that tail is never materialized. materialized_only rollups read it EMPTY under a floor that says covered; real-time baselines serve it from neither branch. The floor-only probes cannot see a hole above the floor.

What lands. RepairMaterializationHolesAsync in a new partial TimescaleSupport.MaterializationHoles.cs: for every registered aggregate in RollupBackfill.Targets order (raw-sourced before hierarchical, so a daily's scan sees its hourly's repair) plus the baselines, read the materialized span off the materialization hypertable, scan from max(floor, source horizon) to the last materialized bucket with two correlated EXISTS probes per bucket (the materialization by bucket, the source by time range with the aggregate's own WHERE, so a restart hour on an interval-honest successor is not a hole), merge contiguous buckets, cap at one refresh policy window per aggregate per start (24 hourly / 3 daily), oldest first, and refresh each range over exactly its bounds. Plain refresh first, forced only on a measured remainder — the backfill's own escalation shape. One INFORMATION line per hole (bounds, buckets, path, duration); deferred ranges named with their bounds; failure isolated per aggregate. Launched by the worker right after the ensure on its own connection, not awaited (the baseline backfill's reason), drained at shutdown. No start_offset changes anywhere.

What the rig corrected. The first cut went straight to force on the reasoning that the tail's rows were inserted above the invalidation threshold and never logged. That is false on TimescaleDB 2.28.1: when a refresh advances the threshold past a region it did not cover, the engine records it as invalid, and the plain refresh materializes the outage tail. The forced path is kept for the hole with no invalidation behind it (materialization rows lost, a refresh cut short) — planted live, shown to survive a plain refresh and materialize on the forced one.

What this does NOT do

  • Does not touch PgMigrations.cs, StorageVersion, MeasureCatalog, DarlingRetention*, or any PgTarget* file. No rung.
  • Does not retire the legacy trio and does not build daily successors (stated why; the daily compression band is full).
  • Does not gate the raw purge on the successors.
  • Does not change any hourly-trend rate's denominator (the successors carry sample_interval_seconds_sum for a future rule-6 reader; today's hourly trend still divides by the bucket width).
  • Does not widen any start_offset.

Tests

  • New IntervalHonestHourlyRollupTests (8 pure + 1 live) and MaterializationHoleRepairTests (7 pure + 1 live). Red-first shown by mutation: short-circuiting HourlyRelationFor and stripping the successor WHERE → 5 reds; removing the forced escalation → the no-invalidation leg red.
  • Live on PG 18.4 / TimescaleDB 2.28.1 (the CI tag): legacy 4 samples / floor 0 vs successor 3 / floor 1,000 / 600 s of interval, sums equal; the store's own probe routes by coverage on both store shapes; the outage sequence planted step for step, scan finds exactly the tail, one refresh into a compressed materialization chunk materializes it, second pass finds nothing, the deleted-materialization hole materializes only on force.
  • 19 pre-existing grid/census pins updated WITH derivation; full Darling.Tests (11,144) live and non-live diffed against a pristine dev baseline on the same rig: no new failures.

…ple: interval-honest successors (sample_interval_seconds IS DISTINCT FROM 0 baked in, the measured interval summed) join the registry beside the legacy trio, every hourly-tier reader takes the successor where it reaches as far as the legacy, and the phase grid re-derives itself by its own method (partial #3653, item 16 / Q12)

query_stats_hourly, procedure_stats_hourly and query_stats_db_hourly aggregate the (0, 0) row a
restart writes: count(*) AS sample_count counts it and min(delta_*) reads it as a real floor. A
continuous aggregate cannot be altered, so the fix is #3698's shape - query_stats_interval_hourly,
procedure_stats_interval_hourly, query_stats_db_interval_hourly, identical dimensions and columns,
the collector's verdict in the WHERE, sum(sample_interval_seconds) carried, WITH NO DATA, on the
--backfill-rollups plan through RollupViews.

Where this departs from #3698, and why: the legacy trio STAYS registered. Each has an indefinite
daily tier hierarchical from it (a CAGG's source is fixed at CREATE), so it cannot be dropped
without cascading the daily nor frozen without stopping it. The successors are therefore APPENDED
to HourlyAggregates and the grid re-derives by the documented method: 16 policies, 15 light,
heaviest at :18, window 18 min = 1,080 s, watch line 900 s - 4 s above the 896 s ceiling, the
"18 whole minutes" the ceiling essay names as the floor. The compression band's minute (:35) and
minutes (:36-:59) do not move; the daily aggregate compression band is now FULL at 23 of 23 hours.
Two consequences are pinned rather than found: the 900 s line coincides with #2136's cadence knob,
and the 952 s clean run RefreshCeilingStalenessTests quotes now classifies ApproachingSlot.

Readers: one supply rule - TimescaleSupport.PrefersSuccessor, the rule PgBaselineProvider applied
per server, now shared - applied to a store's floors by RollupCoverage.HourlyRelationFor AFTER the
tier is decided over the legacy pair. Compose, the MCP duration trends and query history, the
daily summary (calendar + get_daily_health) and the FinOps workload readers all route through it;
an absent successor is never named. LogSupersededHourlyRollupCoverageAsync reports the hand-over
on every start. The raw purge is not gated on the successors (the #1661 precedent), stated.

Census: rule 5's roster keeps its five names - the legacy text still exists - and gains the
invariant that every member is superseded by a registered successor carrying the predicate.

Measured on PG 18.4 / TimescaleDB 2.28.1: legacy 4 samples / floor 0, successor 3 / floor 1,000 /
600 s of interval, sums equal; the store's own probe routes by coverage on both store shapes.
…der a floor that said covered: at service start, every continuous aggregate is scanned for bucket ranges its source holds rows for and it never materialized, and each is closed with one refresh over exactly its bounds - not a wider start_offset (partial #3653, item 9 / Q10)

The hole, read honestly at #3698: a refresh policy re-materializes [now - start_offset, now -
end_offset] and nothing else, so when the service is down longer than the hourly window the first
refresh after resume opens past the raw collected between the last pre-outage refresh's window end
and the moment collection stopped - about two hours - and that tail is never materialized. A
materialized_only rollup reads it EMPTY under a floor that says covered; a real-time baseline
aggregate serves it from neither branch. The floor-only probes (startup backfill, --backfill-rollups)
cannot see a hole above the floor. Ruled Q10: a targeted refresh at start, because a wider
start_offset taxes every hourly refresh forever for a shape that happens once per outage.

RepairMaterializationHolesAsync (a new partial, TimescaleSupport.MaterializationHoles.cs): for every
registered aggregate in RollupBackfill.Targets order (raw-sourced before hierarchical, so a daily's
scan sees its hourly's repair) plus the baselines, read the materialized span off the materialization
hypertable, scan from max(floor, source horizon) to the last materialized bucket with two correlated
EXISTS probes per bucket - the materialization by bucket, the source by time range WITH THE
AGGREGATE'S OWN WHERE, so an hour the aggregate rejects (a restart hour on an interval-honest
successor) is not a hole - merge contiguous buckets, cap at one refresh policy window per aggregate
per start (24 hourly / 3 daily) oldest first, and refresh each range. Plain refresh first, forced
only on a measured remainder: the first cut went straight to force on the reasoning that the tail's
rows were never logged as invalidations, and the rig proved that FALSE on 2.28.1 - the engine records
the region a refresh skips past as invalid, and the plain refresh closes the outage tail. The forced
path is kept for the hole with no invalidation behind it (materialization rows lost, a refresh cut
short), which the rig also plants and which the plain refresh demonstrably leaves standing. One
INFORMATION line per hole with bounds, buckets, path and duration; deferred ranges named; failure
isolated per aggregate. Launched by the worker right after the ensure, not awaited (the baseline
backfill's reason), drained at shutdown.

Measured on PG 18.4 / TimescaleDB 2.28.1: the outage sequence planted step for step, the tail read
empty with raw holding it, the scan found exactly [H3, H5) and not the outage hours, one plain
refresh into a COMPRESSED materialization chunk closed it, a second pass found nothing; H1 deleted
from the materialization survived a plain refresh and closed on the forced one, the pass saying so.
@erikdarlingdata
erikdarlingdata force-pushed the fix/3653-hourly-successors-hole branch from a7ecf42 to 6213508 Compare September 19, 2026 12:23
@erikdarlingdata
erikdarlingdata merged commit a4ca372 into dev Sep 19, 2026
6 of 7 checks passed
@erikdarlingdata
erikdarlingdata deleted the fix/3653-hourly-successors-hole branch September 19, 2026 12:33
This was referenced Sep 19, 2026
erikdarlingdata added a commit that referenced this pull request Sep 20, 2026
…er skipped below its ceiling, and names the day in days_missing[] — the hole scan's two probes per server at day grain, on both daily tools, the viewer's day-detail line and the web calendar (partial #3653 — item 9, A6)

The lie: a day at or below a server's rollup ceiling for which the rollup held no row fell out of the daily
summary's LEFT JOIN and was COALESCEd to unique_queries = 0 beside that day's real wait, CPU and deadlock
numbers — the same 0 a measured-quiet day prints. What happened is that the tier never materialized the day
(the pre-outage tail RetentionTierRouter's essay derives; repaired at the next start by #3731, so rarer, not
gone: between the post-resume refresh that moves the ceiling and that start, and past the per-start cap, the
calendar still read the skipped day).

The witness is TimescaleSupport.MaterializationHoleScanSql's own definition, per server, at day grain — no
second one: the routed queries CTE gains a third UNION ALL member emitting (day, NULL) for every day at or
below the server's ceiling where the rollup holds no row for the server AND the rollup's registered source
(MaterializationHoleTargets: the daily's legacy hourly, the hourly's raw query_stats, with the interval-honest
successor's own WHERE carried) holds an admitted row. A server that ran nothing that day has no source row
and is not named. The outer select projects CASE WHEN q.d IS NULL THEN 0 ELSE q.c END (a no-op on raw) and
the presence arm reads q.c so a not-carried day is not a source that holds the day. Readers keep long?;
DailySummaryRangeReadResult.DaysMissing derives from the NULLs; get_daily_summary and get_daily_summary_range
emit unique_queries: null and a trailing days_missing[]; PerformanceCalendarDay.UniqueQueries and
BuildKeyMetricsLine take long? ("not materialized at this tier"); the web calendar becomes a composite that
renders "not materialized" and one notice line from days_missing. Lite has no rollup tier and is untouched
(the shared model's field widening is a no-op there).

Proven live on TimescaleDB 2.28.1 / PG18: the skipped day reads NULL at the daily tier and is listed, the
empty day is absent, the day past the ceiling reads from raw, the hourly that carried everything has no
NULL, the restart-only day is a hole for the legacy and not for the successor, the reader routes the
100-day window to Daily and the wire carries null + days_missing, and RepairMaterializationHolesAsync then
closes exactly the NULL day. Planner: (server_id, bucket) index probes per day, 2 ms over a 31-day window
with a 41-server population.
erikdarlingdata added a commit that referenced this pull request Sep 20, 2026
…er skipped below its ceiling, and names the day in days_missing[] — the hole scan's two probes per server at day grain, on both daily tools, the viewer's day-detail line and the web calendar (partial #3653 — item 9, A6) (#3788)

The lie: a day at or below a server's rollup ceiling for which the rollup held no row fell out of the daily
summary's LEFT JOIN and was COALESCEd to unique_queries = 0 beside that day's real wait, CPU and deadlock
numbers — the same 0 a measured-quiet day prints. What happened is that the tier never materialized the day
(the pre-outage tail RetentionTierRouter's essay derives; repaired at the next start by #3731, so rarer, not
gone: between the post-resume refresh that moves the ceiling and that start, and past the per-start cap, the
calendar still read the skipped day).

The witness is TimescaleSupport.MaterializationHoleScanSql's own definition, per server, at day grain — no
second one: the routed queries CTE gains a third UNION ALL member emitting (day, NULL) for every day at or
below the server's ceiling where the rollup holds no row for the server AND the rollup's registered source
(MaterializationHoleTargets: the daily's legacy hourly, the hourly's raw query_stats, with the interval-honest
successor's own WHERE carried) holds an admitted row. A server that ran nothing that day has no source row
and is not named. The outer select projects CASE WHEN q.d IS NULL THEN 0 ELSE q.c END (a no-op on raw) and
the presence arm reads q.c so a not-carried day is not a source that holds the day. Readers keep long?;
DailySummaryRangeReadResult.DaysMissing derives from the NULLs; get_daily_summary and get_daily_summary_range
emit unique_queries: null and a trailing days_missing[]; PerformanceCalendarDay.UniqueQueries and
BuildKeyMetricsLine take long? ("not materialized at this tier"); the web calendar becomes a composite that
renders "not materialized" and one notice line from days_missing. Lite has no rollup tier and is untouched
(the shared model's field widening is a no-op there).

Proven live on TimescaleDB 2.28.1 / PG18: the skipped day reads NULL at the daily tier and is listed, the
empty day is absent, the day past the ceiling reads from raw, the hourly that carried everything has no
NULL, the restart-only day is a hole for the legacy and not for the successor, the reader routes the
100-day window to Daily and the wire carries null + days_missing, and RepairMaterializationHolesAsync then
closes exactly the NULL day. Planner: (server_id, bucket) index probes per day, 2 ms over a 31-day window
with a 41-server population.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant