TimescaleSupport.HeaviestHourlyRefreshObservedCeilingSeconds is 594, pinned by RefreshCeilingProvenancePinTests to the maximum of a published 16-run population (194 s–594 s, mean 352.5, median 345) for policy_refresh_continuous_aggregate query_store_stats_interval_hourly (job 1054).
Three runs since then exceed it, read from the store's own job telemetry:
| when |
duration |
vs the pinned ceiling |
| 2026-09-07 00:00 hour |
702 s |
1.18× |
| 2026-09-07 21:00 hour |
871 s |
1.47× |
| 2026-09-08 00:00 hour |
815 s |
1.37× |
get_store_metrics currently reports this job at last_run_duration_ms = 815116, duty_vs_cadence = 22.6%, 375 runs, 2 failures — the only policy on the store with failures other than policy_telemetry.
This is the constant's own pre-registered trigger, not a new opinion
Its summary states the limitation and what would invalidate it:
What would change that answer … or this maximum ceasing to be a lower bound on the truth. The second is the live limitation, and it is a property of the READ rather than of the estimator: the boundary day's runs are a census of that day after the boundary, while the days after it are a SAMPLE …
It also names the census read that settles it rather than leaving it implied: timescaledb_information.job_history, whose succeeded column this store has because DarlingManagedPostgres.BuildConfAppend enables timescaledb.enable_job_execution_logging (#1681) — read as every run since the boundary rather than sampled, giving count, median and maximum each side of it.
The grid is NOT breached, and the margin is what changed
Stating this explicitly because it is the mistake to avoid: RefreshPhaseSlotSeconds is 900 s (RefreshPhaseStepMinutes = 15), and 871 s fits. Nothing has overrun its slot.
What has changed is the clearance. The doc deliberately presents the bound and its margin as two numbers — RefreshPhaseSlotSeconds minus the ceiling — so:
|
seconds |
| slot |
900 |
| ceiling as pinned |
594 |
| clearance today |
306 |
| clearance if the population is re-derived at 871 |
29 |
And 871 s is already above RefreshSlotWarningSeconds (900 × 5/6 = 750 s), which is the product's own signal that a slot is being consumed.
So the ask is not "raise the constant" — it is: re-derive the population by the method the doc records, using the named census read, and then decide whether 15-minute slots still hold. TimescaleSupportTests holds the grid clear of this constant, so a re-derivation to 871 may redden it, which is the check working rather than a problem with the check.
What I am NOT claiming
I found these runtimes while looking for the mechanism behind #3112's lag band, and I got the mechanism wrong twice before the arithmetic corrected me — first as "the policies are not staggered" (they are, by an initial_start phase grid), then as "the grid's ceiling is breached" (it is not; 871 fits in 900). I am not asserting this explains #3112. The band's own scoring is on #3112 and this issue is only the calibration observation, which stands on its own regardless of what #3112 turns out to be.
One adjacent observation, recorded without a claim attached: the hourly compression policies on the same hypertable family have last-run durations of 552 s (query_store_stats), 360 s (query_stats) and 198 s (query_snapshots), all well above OtherHourlyRefreshObservedCeilingSeconds (140 s). Whether a compression run of that length crossing an adjacent phase boundary matters is a question for whoever owns the grid — the doc says non-heaviest slots are "treated as occupied only at their start", and I do not know whether that reasoning was meant to cover compression runtimes of this size.
Related: #3035 (the phase grid), #3101 (deriving a constant as a method), #3012, #3112.
TimescaleSupport.HeaviestHourlyRefreshObservedCeilingSecondsis 594, pinned byRefreshCeilingProvenancePinTeststo the maximum of a published 16-run population (194 s–594 s, mean 352.5, median 345) forpolicy_refresh_continuous_aggregate query_store_stats_interval_hourly(job 1054).Three runs since then exceed it, read from the store's own job telemetry:
get_store_metricscurrently reports this job atlast_run_duration_ms = 815116,duty_vs_cadence = 22.6%, 375 runs, 2 failures — the only policy on the store with failures other thanpolicy_telemetry.This is the constant's own pre-registered trigger, not a new opinion
Its summary states the limitation and what would invalidate it:
It also names the census read that settles it rather than leaving it implied:
timescaledb_information.job_history, whosesucceededcolumn this store has becauseDarlingManagedPostgres.BuildConfAppendenablestimescaledb.enable_job_execution_logging(#1681) — read as every run since the boundary rather than sampled, giving count, median and maximum each side of it.The grid is NOT breached, and the margin is what changed
Stating this explicitly because it is the mistake to avoid:
RefreshPhaseSlotSecondsis 900 s (RefreshPhaseStepMinutes = 15), and 871 s fits. Nothing has overrun its slot.What has changed is the clearance. The doc deliberately presents the bound and its margin as two numbers —
RefreshPhaseSlotSecondsminus the ceiling — so:And 871 s is already above
RefreshSlotWarningSeconds(900 × 5/6 = 750 s), which is the product's own signal that a slot is being consumed.So the ask is not "raise the constant" — it is: re-derive the population by the method the doc records, using the named census read, and then decide whether 15-minute slots still hold.
TimescaleSupportTestsholds the grid clear of this constant, so a re-derivation to 871 may redden it, which is the check working rather than a problem with the check.What I am NOT claiming
I found these runtimes while looking for the mechanism behind #3112's lag band, and I got the mechanism wrong twice before the arithmetic corrected me — first as "the policies are not staggered" (they are, by an
initial_startphase grid), then as "the grid's ceiling is breached" (it is not; 871 fits in 900). I am not asserting this explains #3112. The band's own scoring is on #3112 and this issue is only the calibration observation, which stands on its own regardless of what #3112 turns out to be.One adjacent observation, recorded without a claim attached: the hourly compression policies on the same hypertable family have last-run durations of 552 s (
query_store_stats), 360 s (query_stats) and 198 s (query_snapshots), all well aboveOtherHourlyRefreshObservedCeilingSeconds(140 s). Whether a compression run of that length crossing an adjacent phase boundary matters is a question for whoever owns the grid — the doc says non-heaviest slots are "treated as occupied only at their start", and I do not know whether that reasoning was meant to cover compression runtimes of this size.Related: #3035 (the phase grid), #3101 (deriving a constant as a method), #3012, #3112.