Repository navigation
Raw retention purges run only when the service triggers them, never at PostgreSQL start (#4299) - #4391
Merged
Conversation
Adds the SQL primitives for #4299's (d') design: the three raw retention jobs (query_stats, procedure_stats, query_store_stats) move their armed/held verdict off TimescaleDB's own scheduled flag onto config->>'darling_armed', so the service can trigger the purge itself instead of TimescaleDB's own job runner racing the service at start. - ConvergeRawArmedStateSql: forces scheduled = false unconditionally and merges (||, never a replacing =) darling_armed into config. - RawArmedReadExpression / RawArmedStateSql: the shared read that treats a missing darling_armed key as held, never armed. - Five unit pins in RawArmedStateConvergeTests.cs. Scope: lane 4299-1a only (converge SQL + read helper + unit pins). Does not wire these into EnsureRetentionPoliciesAsync's sweep loop, add the run_job primitive, or touch existing scheduled-based tests -- that is lane 4299-1b, next on this branch.
…e fresh (#4299 L4) Confirmed #4186's seam-window fix (MaterializationHoleScanWindows) already extends the scan's LOWER bound for a stitched successor's un-materialized seam tail, but that is scoped to the successor/legacy seam case, not the general raw-sourced case. The remaining gap: the ORDINARY window's lower bound for a raw-sourced target (query_stats, procedure_stats, query_store_stats) still clamped to utcNow minus MaterializationHoleScanSpanFor, a horizon that assumed raw purged on its own schedule. Under #4299's variant (d prime) the three raw jobs never run on a schedule at all, so raw can legitimately hold rows far older than that clamp while the service- triggered purge has not fired yet, and the old clamp would leave a real hole below it unscanned indefinitely. Fix: IsRawSourced(source) picks out the three RawTierCoverage relations; for those, RawFloorHorizonAsync reads the raw table's own min(time column) as the scan's lower bound instead of the time-based horizon (falling back to the old horizon when raw is empty, where it is harmless). Also builds the function L2's trigger will call as its gate: HoleFreeThroughAsync(connection, target, materialization, dropFrom, dropTo, deferredRanges) -- a FRESH scan of the exact range a purge would drop, never a reuse of a stale repair result, that also treats any deferred range (from the same pass's CapMaterializationHoleRepairs) overlapping the drop window as a hole even though a bare scan of already-repaired buckets would read clean. Pins added to MaterializationHoleRepairTests.cs: - IsRawSourced_IsTrueOnlyForTheThreeRawTierCoverageRelations - DeferredRangeOverlappingDropWindow_WouldBlockTheGate_ByHalfOpenIntervalOverlap Both RED on dev (the members don't exist there -- confirmed by grep). No migrations. Only TimescaleSupport.MaterializationHoles.cs and its test file touched, per brief -- did not touch TimescaleSupport.cs (4299-1a's file).
For the three raw retention jobs (query_stats, procedure_stats, query_store_stats — RawTierCoverage), EnsureRetentionPoliciesAsync now runs ConvergeRawArmedStateSql instead of ArmRetentionPolicySql / HoldRetentionPolicySql: scheduled stays permanently false and the armed/held verdict is written to config->>'darling_armed' instead. Every other retention job keeps today's scheduled semantics unchanged. Because this runs inside the sweep (called at Startup and from the hourly Periodic pass per #3812), the converge now happens on every pass — a DBA's alter_job(scheduled => true) on a raw job gets reverted to scheduled = false with the armed key restored within the hour, or at the next start. The prior-state read (used to tell a transition from a re-assertion, for the Armed/ReHeld log lines and tally) now reads RawArmedStateSql for a raw relation instead of RetentionPolicyScheduledSql, so the transition detection stays correct under the new verdict. Adds RawRelations (RawTierCoverage's relation names) as the lookup the sweep uses to branch. Not in this commit: CALL run_job — the raw jobs never purge yet.
…duled-based pins off query_stats (#4299) The indeterminate-coverage branch previously left everything alone for a raw relation too. Per the ruling (the hourly converge runs every pass, unconditionally), a raw relation's scheduled flag now still converges to false even when coverage could not be measured this pass - only the darling_armed key is left untouched on that path. Test migration (query_stats is now one of the three never-scheduled raw relations, so it can no longer stand in for a scheduled-semantics pin): - TimescaleContinuousAggregateTests: ArmRetentionPolicy_Targets... and HoldRetentionPolicy_IsTheMirrorOfArming... now render against QueryStatsHourlyView (a non-raw relation) instead of query_stats. Assertions kept, not deleted. - RetentionReevaluationTests (RetentionReevaluationLiveTests): added RawArmedAsync (reads config->>'darling_armed' via RawArmedStateSql) alongside the existing ScheduledAsync, and every held/armed assertion on HeldRelation (query_stats) now checks the armed key; scheduled is separately asserted to stay false throughout, proving the raw relation never reaches TimescaleDB's own scheduler. - TheArmAndHoldStatements_StayUnconditional...'s source-order anchor updated for the new ternary read (isRawRelation ? RawArmedStateSql : RetentionPolicyScheduledSql).
Four live xUnit pins in a new RawRetentionServiceTriggeredLiveTests class, proving: 1. A product-created raw retention job converges from the old scheduled=true/no-config shape onto scheduled=false with config->>'darling_armed' recorded, and drop_after survives. 2. A DBA's alter_job(scheduled => true) against an already-converged raw job is reverted by the very next sweep, regardless of the coverage verdict that pass reaches. 3. A non-raw retention relation keeps today's arm/hold semantics (scheduled flag, no darling_armed key). 4. A Covered raw job reads darling_armed = 'true' but scheduled stays false (never true, which is how the pre-fix code armed a raw relation). All four verified GREEN in-process (dotnet Darling.Tests.dll, Release, WindowsDesktop entry stripped) against a real TimescaleDB 2.30.1/PG18 rig, and RED against 8516650 (has the SQL, no wiring).
…s scheduled (#4299) query_store_stats and query_stats are raw relations under #4299's design: their armed verdict lives in config->>'darling_armed', converged through ConvergeRawArmedStateSql/RawArmedStateSql, while their 'scheduled' column is unconditionally driven to false. The test's IsArmedAsync still read only 'scheduled', so it always saw false for these two relations regardless of the real armed state. Mirrors the product's own branch (RawRelations.Contains) so the test and the product read the same column for the same relation. No assertion intent changed.
…arling_armed (#4299) RollupCreatedOverExistingHistory_ReadsEmptyBelowItsFloor_UntilTheBackfillFixesIt and InterruptedBackfill_DoesNotLookComplete_AndTheArmingGateRefusesUntilItGenuinelyIs both asserted query_stats' retention job's 'scheduled' column via a literal count query (ArmedRawPolicySql). query_stats is a raw relation under #4299's design: its armed verdict lives in config->>'darling_armed' via RawArmedStateSql, while 'scheduled' is unconditionally driven to false. Replaced ArmedRawPolicySql's two call sites with a RawArmedAsync helper that reads through RawArmedStateSql. Kept ArmedHourlyPolicySql/its call site unchanged: query_stats_interval_hourly is not a raw relation and still uses scheduled semantics. No assertion intent changed.
…#4299) Pass 1b's hand-arm (the runbook-forbidden alter_job the test uses to set up a re-hold) only wrote scheduled => true. Under #4299 (d') query_stats' real armed state lives in config->>'darling_armed', read via RawArmedStateSql; EnsureRetentionPoliciesAsync's prior-state read for a raw relation consults that key, not scheduled. The hand-arm therefore never registered as armed from the sweep's point of view, so pass 1b's transition to re-held never happened (re-held stayed 0). Fixed the hand-arm to also merge darling_armed=true into the job's config, matching what ConvergeRawArmedStateSql would have written, and added an assertion that it took. Assertion intent unchanged; this is completing the raw-relation migration this test file had already started (RawArmedAsync existed but the setup step feeding it a transition to detect did not).
…heduled (#4299) query_stats and procedure_stats are raw relations under #4299's design: their armed verdict lives in config->>'darling_armed' via RawArmedStateSql, while 'scheduled' is unconditionally driven to false. IsRetentionScheduledAsync only read 'scheduled', so it always saw false for both raw relations in RawPurge_ArmsOffSuccessorHourlyCoverage_NotTheEmptyLegacyOnes regardless of the real armed state. Now branches on TimescaleSupport.RawRelations, mirroring the product's own branch, so the test reads the same column the product reads for the same relation. No assertion intent changed.
…d raw purge (#4299) Lane 4299-1c on the shared branch. - RunRetentionPurgeJobSql/RunRetentionPurgeJobAsync: the service's own CALL run_job(id) trigger for a raw retention job under variant (d'), wrapped in BEGIN; SET LOCAL lock_timeout; CALL run_job($1); COMMIT; with a bounded lock_timeout (RunRetentionPurgeJobLockTimeout, 5s) and a named CommandTimeout (RunRetentionPurgeJobTimeoutSeconds = SetupTimeoutSeconds). A timeout or any other failure is logged at Warning, best-effort ROLLBACK recovers the connection, and the method returns false so the caller's next hourly pass retries. Not called from anywhere yet (L2's job). - Two rig probes, run on real TimescaleDB 2.30.1/PostgreSQL 18 containers, both PASS: 1. lock_timeout applies INSIDE run_job: a second session held ACCESS EXCLUSIVE on the target chunk, CALL run_job wrapped in BEGIN/SET LOCAL/COMMIT raised 55P03 lock_not_available at ~5s (not 30s of blocking), chunk count unchanged after. Also proven as a new xUnit live pin, RawRetentionPurgeJobLiveTests, GREEN in-process on a real rig (0.7s). 2. crash safety: started run_job on a never-scheduled job holding a lock, then ended it two ways in parallel containers - pg_terminate_backend and SIGKILL (forces a postmaster crash-restart). After each: 10 minutes of 30s polling showed no scheduler run of the job (total_runs/last_run_started_at unchanged) and the chunk count held at 40 in both trials. - Docs: RetentionArmSafetySql's XML doc and the "not load-bearing" comment in TimescaleSupport.MaterializationHoles.cs (plus its twin in DarlingWorker.cs) now state that the raw purge only runs when the service triggers it after a Covered verdict and a finished repair under the current pg_postmaster_start_time(), naming both exceptions: the first start after the upgrade, and a DBA's alter_job/run_job before the next converge. Build: Darling.Tests and Lite.Tests both 0/0 warnings/errors with -p:EnableWindowsTargeting=true. StorageCommandTimeoutTests, LiveCleanupConversionRatchetTests, StartupCommandTimeoutTests, DocCommentHygiene, AlertReadFailureSurfaceTests all green in-process.
…ix/4299-service-triggered-raw-purge
TriggerRawPurgeCoreAsync(connection, logger, cancellationToken) takes the exact body of the private instance method, unchanged, so Darling.Tests (already reachable via this project's InternalsVisibleTo) can drive the trigger directly against a live store for the end-to-end pins. The instance method is now a one-line forward. No behaviour change.
…, surface SqlState
RunRetentionPurgeJobSql sent BEGIN; SET LOCAL lock_timeout; CALL run_job($1::integer); COMMIT;
as one NpgsqlCommand with a positional parameter. Npgsql's extended protocol rejects that with
42601 ("cannot insert multiple commands into a prepared statement"), so the CALL never ran and
RunRetentionPurgeJobAsync always returned false, silently. The existing lock-timeout pin only
asserted false + elapsed<20s, so it passed on the 42601 without ever exercising the real path.
Proven on the rig with psql (unblocked): BEGIN; SET LOCAL lock_timeout='5s'; CALL run_job(<id>);
COMMIT; succeeds and drops the chunk. Fixed by opening an explicit NpgsqlTransaction and sending
SET LOCAL and CALL as separate commands on it.
RunRetentionPurgeJobAsync now returns a RetentionPurgeOutcome(bool Ran, string? SqlState) instead
of a bare bool, so a caught failure states which SQLSTATE it was. TriggerRawPurgeCoreAsync (the
one caller) logs it.
Pins (live, Darling/Darling.Tests/RawRetentionPurgeJobLiveTests.cs):
- RunRetentionPurgeJob_Unblocked_RunsAndDropsTheOldChunk (new): unblocked with a chunk older than
drop_after, asserts Ran==true and the chunk count drops. RED on 0c3ed6a: false, count unchanged.
- RunRetentionPurgeJob_BlockedOnAConflictingLock_...: now also asserts SqlState=="55P03", so it
can never pass on a 42601 again. RED on 0c3ed6a: SqlState was 42601, not surfaced by the old
bare-bool return.
Both pins run GREEN in-process on this branch (Darling.Tests.dll, macOS, TimescaleDB 2.30.1).
Copies the lane-4299-2d draft RawPurgeTriggerLiveTests.cs into the repo. All five [Fact]s pass in-process against the fixed RunRetentionPurgeJobAsync from 9329306: no product or seed changes were needed.
…tency Extracts DarlingWorker.ShouldLaunchMaterializationHoleRepair as a pure internal static method (behaviour unchanged) so a second service's relaunch decision can be asserted directly. Adds RawPurgeTriggerLiveTests.TwoServices_RelaunchGuardHolds_PurgeRunsOnceOnly: two independent NpgsqlDataSources against the same scratch store prove Service B sees Service A's epoch stamp, decides not to launch a repair, re-stamps 0 rows, and that TriggerRawPurgeCoreAsync called from both services in turn drops chunks once and is a no-op the second time.
…r raw jobs - RetentionHoldReadSql now projects config->>'darling_armed' as an extra column; ReadRetentionHoldReadingsAsync picks darling_armed for one of the three raw relations (TimescaleSupport.RawRelations) and keeps j.scheduled for every other retention policy, so the Retention Held alert judges a raw job by its real verdict instead of its permanent scheduled=false. - JobCadenceReadSql excludes the three raw retention jobs from the #2136 cadence reading (the minimum-safe exclusion; cadence-from-trigger is a follow-up lane's work). - ApplyPolicyJobsStuckAsync's not-scheduled WARNING now says a raw job is unscheduled by design (#4299), not a #1680/#1877 coverage hold, when the hypertable is one of the three raw relations.
…text branch - RetentionHoldAndCadenceRawRedirectLiveTests (live, own-store): darling_armed true/false on RetentionHoldReadSql, and JobCadenceReadSql's absence of a raw job even after a successful triggered run. - DarlingSelfAlertTests: the M4 pin (a raw job flagged stuck is never re-armed, GREEN on both commits) and the WARNING-text pin (unscheduled-by-design for a raw relation, RED before this lane's branch).
…eReadSql built from RawRelations
…st_purge alongside the retention-hold pass
…, not the excluded job_stats row
…solvable successor (#4391) The service-triggered raw purge (#4299) read the coverage sweep's standing darling_armed flag to decide whether to drop chunks. That flag is left unchanged whenever a probe cannot measure coverage this pass, so a purge could run on a stale Covered verdict from an earlier hour even though this pass's own coverage read came back Unknown. Separately, when a successor's continuous aggregate could not be resolved (mid-rebuild, renamed, or dropped), the hole check skipped it and treated the raw relation as hole-free. The trigger now measures coverage fresh in the same pass via IsRawTierDropSafeAsync, which answers true only on a Covered verdict from this call; darling_armed keeps its existing meaning for readers and alerts and is untouched. An unresolvable successor now blocks the purge and records the outcome as gate_unknown instead of being skipped. Added two live pins to RawPurgeTriggerLiveTests: a stale darling_armed=true with coverage measured Unknown this pass, and a Covered/current-epoch state whose hole-check successor cannot be resolved. Both drop no chunks on the fixed code.
…nreadable one as never-written (#4299, #4391) - TimescaleSupport.cs: add ReadRawLastPurgeStateAsync, returning a tri-state RawLastPurgeReadState (Present/NeverWritten/ReadFailed) alongside the record. ReadRawLastPurgeOutcomeAsync becomes a thin wrapper so its seven existing callers keep compiling unchanged. RawPurgeOverHorizonReading gains a trailing LastPurgeReadFailed flag (default false, so every existing constructor call still compiles). - DarlingSelfAlertEvaluator.cs: the Raw Purge Over Horizon alert now also fires when the last-purge record is more than two hours old (twice the hourly tick the purge trigger rides on), even if its outcome was "ran" -- an old ran record must not keep the alert quiet forever once the trigger itself has stopped running. A read failure is counted (via the existing swallowed-read counter) and skipped: it neither fires nor resolves the alert, leaving the standing state exactly where it was. Reason text gains gate_unknown and gate_error arms. - DarlingSelfAlertTests.cs: pins for the stale-record fire (with its detail text), the recent-record non-fire, a read failure while the alert is active (must not resolve, must count), a read failure while inactive (must not fire), and the two new reason-text arms.
…he raw purge trigger's own gate failure (#4299, #4391) - RepairEpochStampAllowed(summary) allows the stamp only when the materialization-hole repair pass had zero isolated failures; deferred/remaining holes do not block it (the trigger's own fresh hole scan covers those). - TriggerRawPurgeCoreAsync's per-relation catch now logs at Warning and records the outcome as gate_error (with SqlState when available) instead of silently logging at Information with nothing recorded. - Adds RawPurgeTriggerGateErrorTests.cs: a pure pin for RepairEpochStampAllowed and a live pin that forces the trigger's hole scan to fail with a lock timeout and asserts no chunk drops and the outcome records gate_error/57014.
…riodic repair (#4299, #4391) - ConvergeRawArmedStateSql now writes only when the job is scheduled or its darling_armed key differs from the verdict, matching RawRepairEpochStampSql's guard style; the Unknown branch's schedule-only revert is guarded the same way, and logs one Information line when a DBA's alter_job(scheduled => true) is reverted. - The Periodic relaunch now reads all three raw jobs' repair epoch and relaunches when ANY is stale, not just the first; the log line names the stale relations. - The Periodic pass's own repair launch is kept in a field and awaited in the same shutdown drain as the start-path launch, so it is neither lost nor unobserved on shutdown. - RecordRawLastPurgeOutcomeAsync's record is now built server-side with jsonb_build_object and explicit casts, never string-interpolated.
…urge read (#4299, #4391) RawPurgeConvergeLiveTests.cs: scheduled=false holds after Covered/Short/Unknown when a DBA re-arms a raw job by hand; a second converge pass on an already-converged store writes nothing and logs the revert line exactly once, on the pass that actually reverted something; the tri-state last-purge read distinguishes NeverWritten, ReadFailed (a malformed value) and Present (a real trigger pass) live.
erikdarlingdata
marked this pull request as ready for review
September 26, 2026 14:25
This was referenced Sep 26, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #4299.
Why
An armed raw-retention job could run on TimescaleDB's own schedule when PostgreSQL started, before the service's start sweep could hold it. After a long enough outage, that drops raw history no rollup has materialized yet. The restart probe below shows it: with the old shape, 6 of 10 chunks were gone about 5 seconds after PostgreSQL was ready.
What changes
query_stats,procedure_stats,query_store_stats).config->>'darling_armed'. A missing key reads as held.scheduledflag or armed state actually differs. A DBA'salter_job(scheduled => true)is reverted and logged at Information. That's a converge, not a re-hold: it never arms the job (Held retention policies only arm on service restart - re-evaluate them on a periodic cadence too #3812).scheduledsemantics.TimescaleSupport.RunRetentionPurgeJobAsync:CALL run_job(id)in its own transaction,lock_timeout5 s), and only from the hourly Periodic pass (DarlingWorker.TriggerRawPurgeCoreAsync). It needs, all checked in the same pass:IsRawTierDropSafeAsync), never the storeddarling_armed, which keeps a prior pass'struethrough an Unknown probe (Retention gate is arm-only, so a policy whose coverage list GROWS is never re-held — a store upgrading into a new consumer keeps purging its source #1877). Short, Unknown or a probe error does not purge;pg_postmaster_start_time(), stamped on the jobs asdarling_repair_epoch(microseconds,IS DISTINCT FROM-guarded). A repair with an isolated per-aggregate failure leaves the stamp unwritten;HoleFreeThroughAsync). A successor rollup that can't be resolved (mid-rebuild, renamed, dropped) holds the purge instead of reading as hole-free. The hole scan for raw-sourced rollups now starts at raw's actual floor. Deferred ranges are passed empty: a deferred range is still a hole, so the fresh scan finds it.gate_errorwith its SqlState.darling_armedfor raw jobs;config->>'darling_last_purge'(built server-side withjsonb_build_object);RetentionArmSafetySqland the "ordering is not load-bearing" lines now describe the service-triggered purge and its two exceptions:alter_job/run_jobbefore the next converge.Test plan
Live pins that FAILED on the pre-fix code (behaviour, not only a compile failure):
RawRetentionServiceTriggeredLiveTests: 3 of 4 failed on the pre-fix commit (converge not wired). The fourth can't compile there.RetentionHoldAndCadenceRawRedirectLiveTests, thedarling_armed = truecase: readArmed = falseon an earlier commit.DarlingSelfAlertTests, the raw stuck-job warning text: failed on an earlier commit.RawRetentionPurgeJobLiveTests: before the fix, the purge primitive returned SqlState 42601 on every call (Npgsql's extended protocol rejects a multi-statementBEGIN; SET LOCAL ...; CALL ...; COMMIT;in one command), so the service-triggered purge could never run. The lock pin now requires55P03, and a new success pin requires the chunks to drop.New members, so compile-RED on the pre-fix code:
RawPurgeTriggerLiveTests(6: drop only when Covered + current epoch + no hole; Short, stale epoch, a hole and the Startup pass all drop nothing; a second service's purge drops nothing more and records its own pass), each also asserting the recorded outcome;RawRepairEpochStampTests/RawRepairEpochTriggerLiveTests;RawArmedStateConvergeTests;DarlingSelfAlertTests.Pre-existing tests migrated deliberately from
scheduledtodarling_armed, with no assertion removed:QueryStoreCorrectedRollupLiveTests,RollupBackfillLiveTests,RetentionReevaluationLiveTests,FrozenRollupLiveTestsandTimescaleContinuousAggregateTests. Windows CI was green after the migration.Added after review, each run in-process on macOS:
RawPurgeTriggerLiveTests): a staledarling_armed = truewith an Unknown probe, and an unresolvable successor. Both FAILED at runtime on the pre-fix commit (the chunk count went 10 → 4). Both drop nothing now, recordingnot_covered/gate_unknown.RawPurgeTriggerGateErrorTests: a hole scan made to fail (a held lock +statement_timeout) drops nothing and recordsgate_errorwith57014; the clean-repair stamp rule.RawPurgeConvergeLiveTests:scheduled = falseafter every verdict branch (Covered, Short, Unknown);DarlingSelfAlertTests: a stale record fires; a recent one doesn't; an unreadable record neither fires nor resolves and is counted; the new reason texts.Class totals at
cac3ede73:RawPurgeTriggerLiveTests8/8;RawPurgeTriggerGateErrorTests2/2;RawPurgeConvergeLiveTests5/5;DarlingSelfAlertTests248/248;AlertReadFailureSurfaceTests27/27;RawRetentionPurgeJobLiveTests2/2;RetentionHoldAndCadenceRawRedirectLiveTests3/3;FrozenRollupLiveTests21/21;MaterializationHoleRepairTests20/20;TimescaleContinuousAggregateTests44/44 (after merging dev's Hold the raw purge on any hole in the frozen legacy's span, and fill the successor down to raw's floor (#4301) #4401/Don't re-hold the raw purge for an empty successor mid-upgrade (#4300) #4406).The commits after
cac3ede73change comment text, one log level (Debug for the expected raw stuck-job line, now pinned at Debug), three test seed names and one doc-block placement. CI is green ond3d7e252a: every Darling build and PostgreSQL test job, and all Lite jobs.d3d7e252a.Notes for review
AlertReadFailureSurfaceTests: the new alert's wrapper catch is added tos_exemptions(it matches its StaleMute/WebTls siblings, which get their evidence as parameters).DarlingSelfAlertEvaluator.csgoes(…, 12, 14)→(…, 12, 15), and the total exempt count 30 → 31. The alert's record read is counted separately, as a read failure.RetentionReevaluationLiveTests' hand-arm step writes bothdarling_armedandscheduled, which is how a real re-hold is exercised. A DBA's barescheduled => trueis a converge (logged), not a re-hold.Findings during development
RawTierCoverageinTimescaleSupport.csnames exactly three raw relations —query_stats,procedure_stats, andquery_store_stats.lock_timeoutapply insiderun_job? Yes. With a second session holdingACCESS EXCLUSIVEon the oldest chunk for 30s,CALL run_job($1)wrapped inBEGIN; SET LOCAL lock_timeout='5s'; ...; COMMIT;raisedERROR: canceling statement due to lock timeoutat 5.07s elapsed, not after the full 30s hold. Chunk count was unchanged after the failed call.CALL run_jobon a held lock, then ending the blocked backend (pg_terminate_backendin one trial,SIGKILLin another, with a confirmed crash-restart log line in the SIGKILL trial), leftjob_stats.total_runs/last_run_started_atunset and the chunk count unchanged over a 10-minute, 20-poll window in both trials. The next hourly pass is the retry path; nothing needed manual cleanup.RunRetentionPurgeJobSqlsentBEGIN; SET LOCAL lock_timeout = '5s'; CALL run_job($1::integer); COMMIT;as oneNpgsqlCommandwith a positional parameter. Npgsql's extended (prepared-statement) protocol rejects multiple statements in one command (SqlState 42601). The existing lock-timeout pin only assertedfalse+ elapsed<20s, which the 42601 also satisfies, so nothing had shown the success path. Fixed by using an explicitNpgsqlTransactionwithSET LOCAL lock_timeoutandCALL run_job($1::integer)as separate commands, thenCommitAsync.RunRetentionPurgeJobAsyncnow returns aRetentionPurgeOutcome(bool Ran, string? SqlState)instead of a bare bool, so a caller can see the SqlState when a run doesn't happen.CHANGELOG entry
SECTION: Fixed
ENTRY:
query_stats,procedure_statsorquery_store_statsretention job the moment PostgreSQL came up, dropping raw history that no rollup had materialized yet. These three jobs are now never scheduled on TimescaleDB's runner. The service runs the purge itself, from its hourly evaluation only, and only after it measures coverage itself, confirms the hole repair finished cleanly since the current PostgreSQL start, and finds no hole in the range it would drop. A new Raw Purge Over Horizon alert says why a purge didn't run, or that the purge trigger has stopped. The first start after the upgrade still runs the old armed job once.REF:
[Raw retention purges run only when the service triggers them, never at PostgreSQL start (#4299) #4391]: Raw retention purges run only when the service triggers them, never at PostgreSQL start (#4299) #4391