You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
The light ceiling's population floor reads a prose count, but its own published 95th percentile forces a population of 60 #3200
RefreshCeilingProvenancePinTests.PopulationFloor (#3195) is applied at two sites that are not equally guarded, and #3195's description states that asymmetry as an inherent bound rather than a deferred one:
HeaviestHourlyRefreshObservedCeilingSeconds — the floor reads a parsedpopulation.Length over readings the summary actually lists. It cannot be satisfied by a claim.
OtherHourlyRefreshObservedCeilingSeconds — the floor reads the census sentence's count, over a population the summary deliberately does not republish ("a list that long stops being re-derivable by reading and starts being a wall of digits"). A re-derivation stating over 900 runs while having read four satisfies it. The only corroboration is the exclusion-control sentence, which states the same figure a second time — two prose claims agreeing with each other.
That gap can be closed at the light site with figures the file already parses, and without republishing anything. The count stops being an unconstrained claim and becomes one leg of an arithmetic relationship between three independently pinned figures.
The derivation
The light summary publishes, in three different sentences, all three already pinned and all three already read into Verify:
figure
value
pin
census count
874 runs
the light-refresh census, group 3 → lightRuns
95th percentile
42.2 s
the light-refresh quantiles, groups 1–2 → light95Tenths
third-largest reading
140.1 s
the third-largest light run → thirdLargest
The summary also publishes the three largest readings — 226.8, 160.4, 140.1 — and Verify already pins their ordering (midnightMax == lightMax, previousNight < lightMax, thirdLargest < previousNight).
The stated 95th percentile is 42.2 s, below all three. Under nearest rank that is not free — it forces a minimum population size:
the p95 value sits at 1-based index r = ceil(0.95n), which is what NearestRank computes;
three published readings are strictly above that value, so at least three members sit above index r, i.e. n − r ≥ 3;
ceil(0.95n) ≤ n − 3 ⟹ 0.95n ≤ n − 3 ⟹ n ≥ 60.
Checked against NearestRank's own integer form: at n = 59 the rank is 57 and 59 − 57 = 2; at n = 60 the rank is 57 and 60 − 57 = 3. So 60 is the first admissible size, and at the real 874 the rank is 831 with 43 readings above it — comfortably consistent.
So the light site's floor is 60, not 20, and the third of that is bought with no new extraction.
Why this closes the gap rather than restating it
The count is no longer corroborated by a second sentence stating the same number. It is constrained by a relationship among three figures pinned from three different sentences, and to pass, a thin census has to fabricate all three consistently rather than one. A re-derivation that honestly read four runs cannot state a 42.2 s 95th percentile at all: at n = 4 the nearest rank is 4, so the p95 is the maximum.
Concretely, #3195's own successful mutation is caught by this. That mutation cut the census from 874 to 4, rewriting the count sentence and the filter-control sentence consistently, and was green across the whole file before #3195 and green on the quantile clauses after it — the quantile and third-largest sentences were never touched. Under this check, 42.2 < 140.1 demands runs ≥ 60, and 4 is not. That claim is verifiable by re-running exactly that mutation, and it should be verified rather than taken.
It generalises PopulationFloor rather than competing with it
Both are the same search with a different threshold on n − NearestRank(1..n, 95):
≥ 1 — the p95 is a different reading from the maximum → 20, which is PopulationFloor exactly;
≥ 3 — three published readings sit above the p95 → 60.
So derive it the same way PopulationFloor is derived — searched with NearestRank itself, bounded, never written as a literal — and drive the threshold from how many published tail readings actually exceed the stated 95th percentile, counted rather than assumed to be three. If a future re-derivation raised the p95 above the third-largest, the count drops and the bound weakens correctly, because fewer readings above the percentile really is less information about n. PopulationFloor stays as the floor under it: at a count of zero the derived bound degenerates to 1 and the 20-draw floor is the only thing holding the line. Neither weaken nor remove PopulationFloor.
The store-backed route does not work, for two independent reasons
OtherHourlyRefreshRuntimesSql (#3182) is the read that can attribute a live run to one of the twelve non-heaviest hourly views, and that is exactly what it does — a run, singular. It joins timescaledb_information.job_stats USING (job_id) and projects js.last_run_duration, and job_stats carries one row per job. So it returns at most twelve rows, each that view's most recent successful run. It cannot count runs and it cannot measure a span. The premise fails at the instrument, one level above the GUC question.
The read that can count is timescaledb_information.job_history, which is what the census itself used — one row per run. Two things stop it corroborating this count:
The population mismatch, which is fatal on its own. This census is scoped by a positional exclusion rule to runs starting after the 13:44:23 narrowing boundary, and job_history records only from whenever logging was switched on. A live count is therefore a count over a different population, and the summary says so in its own words: "neither is comparable with a sweep over a store's whole job history, because the boundary is there because runs before it worked an unnarrowed HourlyRefreshStartOffset." Comparing them is the cross-regime comparison the constant's own doc warns about, and its stated direction is that a boundary-scoped constant reads as understated by a factor.
So a store read cannot make this count a measurement. What can is arithmetic over what the summary already publishes — which needs no store, no GUC, and no honest-empty handling.
Scope
Raise the light site's effective floor to the derived value, from figures already in hand at that point in Verify.
Keep PopulationFloor and its minimality assertion untouched.
Do not republish the 874 readings — that contradicts the summary's own stated reasoning and trades this weakness for a worse one.
Fixed by #3205, merged to dev as cb208f6a4c05. Closing by hand — closing keywords resolve against main and this landed on dev.
The light ceiling's population size is now constrained by arithmetic over three separately pinned figures instead of by a count stated in prose. DeriveFloorForReadingsAboveThePercentile(k) searches the shipped NearestRank for the first n holding k members above its own 95th percentile: k=0→1, k=1→20 (which is PopulationFloor exactly), k=2→40, k=3→60. The summary states a 42.2 s 95th percentile with a tail of 226.8 / 160.4 / 140.1 — three readings above it — so its own figures force a population of at least 60, three times what #3195's floor required. No literal 20 or 60 appears anywhere, nothing was republished, and TimescaleSupport.cs is unmodified.
It is a second, independent search rather than DerivePopulationFloor delegating to it. The two must agree at k=1 and a test asserts they do; a shared implementation would move both together and the agreement would assert nothing.
Two corrections to this issue as I filed it, both from the lane checking rather than accepting:
The worked example in the issue attributes the failure to the wrong clause.#3195's 874 → 4 mutation does go red, but at the pre-existing 20-draw floor — Verify is fail-fast and 4 is under both bounds. The arithmetic in this issue is right; the attribution isn't. The new bound's actual territory is 20 ≤ n < 60, and that is where it earns its place: this summary already floats "29 runs" as the figure for the current layout, which clears #3195's floor and fails this one.
The paragraph I said to fix is not in the tree. The "cannot be closed here" sentence lives only in #3195's PR description; grep -rn "cannot be closed" over the repo returns nothing. What was genuinely stale is the light site's own comment claiming every figure the paragraph states "is therefore consistent with any count whatever" — both halves of that are now false, and it has been rewritten in the present tense.
The bound is proved exact rather than merely sufficient: consistent census rewrites to 40 and 59 both go red at PIN PERCENTILE TAIL BOUND, and 60 goes green. Instrument mutations too — call deleted, threshold hard-coded to 1, search weakened, the tail list shortened, the filter inverted, all red.
On the choice of where the tail count comes from: it reads the three pins Verify has already required to be strictly ordered, not a sweep of the doc run. The sweep is more general and that is the objection — this run also states a percentile gap, a coverage margin, a growth percentage and a different population's maximum, all above 42.2 s and none of them members, so a sweep inflates the bound in the direction that grants strength. The cost is a frozen list of three whose staleness understates rather than falsely passes.
One thing deliberately not filed: the light doc run states 874 twice more in prose no pin reads. It reads load-bearing and nothing checks it — a narrower question than this one.
RefreshCeilingProvenancePinTests.PopulationFloor(#3195) is applied at two sites that are not equally guarded, and #3195's description states that asymmetry as an inherent bound rather than a deferred one:HeaviestHourlyRefreshObservedCeilingSeconds— the floor reads a parsedpopulation.Lengthover readings the summary actually lists. It cannot be satisfied by a claim.OtherHourlyRefreshObservedCeilingSeconds— the floor reads the census sentence's count, over a population the summary deliberately does not republish ("a list that long stops being re-derivable by reading and starts being a wall of digits"). A re-derivation statingover 900 runswhile having read four satisfies it. The only corroboration is the exclusion-control sentence, which states the same figure a second time — two prose claims agreeing with each other.That gap can be closed at the light site with figures the file already parses, and without republishing anything. The count stops being an unconstrained claim and becomes one leg of an arithmetic relationship between three independently pinned figures.
The derivation
The light summary publishes, in three different sentences, all three already pinned and all three already read into
Verify:the light-refresh census, group 3 →lightRunsthe light-refresh quantiles, groups 1–2 →light95Tenthsthe third-largest light run→thirdLargestThe summary also publishes the three largest readings — 226.8, 160.4, 140.1 — and
Verifyalready pins their ordering (midnightMax == lightMax,previousNight < lightMax,thirdLargest < previousNight).The stated 95th percentile is 42.2 s, below all three. Under nearest rank that is not free — it forces a minimum population size:
r = ceil(0.95n), which is whatNearestRankcomputes;r, i.e.n − r ≥ 3;ceil(0.95n) ≤ n − 3⟹0.95n ≤ n − 3⟹n ≥ 60.Checked against
NearestRank's own integer form: atn = 59the rank is 57 and59 − 57 = 2; atn = 60the rank is 57 and60 − 57 = 3. So 60 is the first admissible size, and at the real 874 the rank is 831 with 43 readings above it — comfortably consistent.So the light site's floor is 60, not 20, and the third of that is bought with no new extraction.
Why this closes the gap rather than restating it
The count is no longer corroborated by a second sentence stating the same number. It is constrained by a relationship among three figures pinned from three different sentences, and to pass, a thin census has to fabricate all three consistently rather than one. A re-derivation that honestly read four runs cannot state a 42.2 s 95th percentile at all: at
n = 4the nearest rank is 4, so the p95 is the maximum.Concretely, #3195's own successful mutation is caught by this. That mutation cut the census from 874 to 4, rewriting the count sentence and the filter-control sentence consistently, and was green across the whole file before #3195 and green on the quantile clauses after it — the quantile and third-largest sentences were never touched. Under this check,
42.2 < 140.1demandsruns ≥ 60, and 4 is not. That claim is verifiable by re-running exactly that mutation, and it should be verified rather than taken.It generalises
PopulationFloorrather than competing with itBoth are the same search with a different threshold on
n − NearestRank(1..n, 95):≥ 1— the p95 is a different reading from the maximum → 20, which isPopulationFloorexactly;≥ 3— three published readings sit above the p95 → 60.So derive it the same way
PopulationFlooris derived — searched withNearestRankitself, bounded, never written as a literal — and drive the threshold from how many published tail readings actually exceed the stated 95th percentile, counted rather than assumed to be three. If a future re-derivation raised the p95 above the third-largest, the count drops and the bound weakens correctly, because fewer readings above the percentile really is less information aboutn.PopulationFloorstays as the floor under it: at a count of zero the derived bound degenerates to 1 and the 20-draw floor is the only thing holding the line. Neither weaken nor removePopulationFloor.The store-backed route does not work, for two independent reasons
OtherHourlyRefreshRuntimesSql(#3182) is the read that can attribute a live run to one of the twelve non-heaviest hourly views, and that is exactly what it does — a run, singular. It joinstimescaledb_information.job_stats USING (job_id)and projectsjs.last_run_duration, andjob_statscarries one row per job. So it returns at most twelve rows, each that view's most recent successful run. It cannot count runs and it cannot measure a span. The premise fails at the instrument, one level above the GUC question.The read that can count is
timescaledb_information.job_history, which is what the census itself used — one row per run. Two things stop it corroborating this count:timescaledb.enable_job_execution_loggingis ON; that GUC defaulted off and could not be healed onto a pre-existing cluster until job_history is silently empty on clusters predating ~2026-08-17, because its GUC lives in the one postgresql.conf block EnsureConfAppended cannot heal #3175/Give timescaledb.enable_job_execution_logging its own conf marker, so existing stores can heal (#3175) #3177 gave it a marker; and a count or a maximum over it on an unhealed store returns zero rows, which reads as "nothing exceeded the line". Any check would have to report the GUC's effective value and source distinctly from the row count —DarlingStoreMetricsReader.GetJobExecutionLoggingAsyncandDarlingMcpStoreMetricsTools.JobHistoryNotealready do exactly that, so the machinery exists.13:44:23narrowing boundary, andjob_historyrecords only from whenever logging was switched on. A live count is therefore a count over a different population, and the summary says so in its own words: "neither is comparable with a sweep over a store's whole job history, because the boundary is there because runs before it worked an unnarrowedHourlyRefreshStartOffset." Comparing them is the cross-regime comparison the constant's own doc warns about, and its stated direction is that a boundary-scoped constant reads as understated by a factor.So a store read cannot make this count a measurement. What can is arithmetic over what the summary already publishes — which needs no store, no GUC, and no honest-empty handling.
Scope
Verify.PopulationFloorand its minimality assertion untouched.