Skip to content

Raw retention purges run only when the service triggers them, never at PostgreSQL start (#4299) - #4391

Merged
erikdarlingdata merged 45 commits into
devfrom
fix/4299-service-triggered-raw-purge
Sep 26, 2026
Merged

erikdarlingdata merged 45 commits into
devfrom
fix/4299-service-triggered-raw-purge

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Refs #4299.

Why

An armed raw-retention job could run on TimescaleDB's own schedule when PostgreSQL started, before the service's start sweep could hold it. After a long enough outage, that drops raw history no rollup has materialized yet. The restart probe below shows it: with the old shape, 6 of 10 chunks were gone about 5 seconds after PostgreSQL was ready.

What changes

  • The three raw retention jobs are never scheduled on TimescaleDB's runner (query_stats, procedure_stats, query_store_stats).
  • The service runs the purge itself (TimescaleSupport.RunRetentionPurgeJobAsync: CALL run_job(id) in its own transaction, lock_timeout 5 s), and only from the hourly Periodic pass (DarlingWorker.TriggerRawPurgeCoreAsync). It needs, all checked in the same pass:
    • a Covered verdict measured by the trigger itself (IsRawTierDropSafeAsync), never the stored darling_armed, which keeps a prior pass's true through an Unknown probe (Retention gate is arm-only, so a policy whose coverage list GROWS is never re-held — a store upgrading into a new consumer keeps purging its source #1877). Short, Unknown or a probe error does not purge;
    • a hole repair that finished with no failures under the current pg_postmaster_start_time(), stamped on the jobs as darling_repair_epoch (microseconds, IS DISTINCT FROM-guarded). A repair with an isolated per-aggregate failure leaves the stamp unwritten;
    • no hole in the range the run would drop, from a fresh scan (HoleFreeThroughAsync). A successor rollup that can't be resolved (mid-rebuild, renamed, dropped) holds the purge instead of reading as hole-free. The hole scan for raw-sourced rollups now starts at raw's actual floor. Deferred ranges are passed empty: a deferred range is still a hole, so the fresh scan finds it.
    • An error inside the trigger's own checks logs at Warning and records gate_error with its SqlState.
  • The Startup pass never purges. A new PostgreSQL start, when no repair is running in the process, launches one repair, and a second service honours the stamp. The Periodic pass relaunches the repair when ANY raw job's epoch is stale, and that repair is awaited at shutdown like the start-path one.
  • Readers and alerts follow the verdict:
    • Retention Held reads darling_armed for raw jobs;
    • the raw jobs leave the Query Store aggregate tax scales serially with fleet size — measure, alert, and evaluate before large onboardings #2136 cadence query and take their cadence from the trigger: the last run's time against the hourly pass;
    • a new "Raw Purge Over Horizon" alert fires for a raw table over its horizon whatever the verdict, and says why the purge didn't run: not covered, repair pending, a hole, an unresolvable successor, a trigger error, or failed with its SqlState. Each attempt is recorded in config->>'darling_last_purge' (built server-side with jsonb_build_object);
    • it also fires when that record is more than 2 hours old (twice the hourly pass), meaning the trigger itself has stopped running. A record that can't be read neither fires nor clears the alert; it's counted as a read failure;
    • the stuck re-arm never touches a raw job (pinned).
  • Docs: RetentionArmSafetySql and the "ordering is not load-bearing" lines now describe the service-triggered purge and its two exceptions:
    • the first start after the upgrade;
    • a DBA's alter_job/run_job before the next converge.
  • No migration.

Test plan

Live pins that FAILED on the pre-fix code (behaviour, not only a compile failure):

  • The restart probe (rig): old shape, 10 → 4 chunks about 5 s after ready; this branch, 0 drops and 0 job runs over 2 minutes (12 polls).
  • RawRetentionServiceTriggeredLiveTests: 3 of 4 failed on the pre-fix commit (converge not wired). The fourth can't compile there.
  • RetentionHoldAndCadenceRawRedirectLiveTests, the darling_armed = true case: read Armed = false on an earlier commit.
  • DarlingSelfAlertTests, the raw stuck-job warning text: failed on an earlier commit.
  • RawRetentionPurgeJobLiveTests: before the fix, the purge primitive returned SqlState 42601 on every call (Npgsql's extended protocol rejects a multi-statement BEGIN; SET LOCAL ...; CALL ...; COMMIT; in one command), so the service-triggered purge could never run. The lock pin now requires 55P03, and a new success pin requires the chunks to drop.

New members, so compile-RED on the pre-fix code:

  • RawPurgeTriggerLiveTests (6: drop only when Covered + current epoch + no hole; Short, stale epoch, a hole and the Startup pass all drop nothing; a second service's purge drops nothing more and records its own pass), each also asserting the recorded outcome;
  • RawRepairEpochStampTests/RawRepairEpochTriggerLiveTests;
  • RawArmedStateConvergeTests;
  • the Raw Purge Over Horizon and raw cadence unit pins in DarlingSelfAlertTests.

Pre-existing tests migrated deliberately from scheduled to darling_armed, with no assertion removed: QueryStoreCorrectedRollupLiveTests, RollupBackfillLiveTests, RetentionReevaluationLiveTests, FrozenRollupLiveTests and TimescaleContinuousAggregateTests. Windows CI was green after the migration.

Added after review, each run in-process on macOS:

  • The fresh-measurement pins (RawPurgeTriggerLiveTests): a stale darling_armed = true with an Unknown probe, and an unresolvable successor. Both FAILED at runtime on the pre-fix commit (the chunk count went 10 → 4). Both drop nothing now, recording not_covered / gate_unknown.
  • RawPurgeTriggerGateErrorTests: a hole scan made to fail (a held lock + statement_timeout) drops nothing and records gate_error with 57014; the clean-repair stamp rule.
  • RawPurgeConvergeLiveTests:
    • scheduled = false after every verdict branch (Covered, Short, Unknown);
    • a second converge on a converged store writes nothing, and the revert line logs once;
    • the three-state record read (never written, unreadable, present).
  • DarlingSelfAlertTests: a stale record fires; a recent one doesn't; an unreadable record neither fires nor resolves and is counted; the new reason texts.

Class totals at cac3ede73:

The commits after cac3ede73 change comment text, one log level (Debug for the expected raw stuck-job line, now pinned at Debug), three test seed names and one doc-block placement. CI is green on d3d7e252a: every Darling build and PostgreSQL test job, and all Lite jobs.

  • The full Windows suite (CI), green on d3d7e252a.

Notes for review

  • AlertReadFailureSurfaceTests: the new alert's wrapper catch is added to s_exemptions (it matches its StaleMute/WebTls siblings, which get their evidence as parameters). DarlingSelfAlertEvaluator.cs goes (…, 12, 14) → (…, 12, 15), and the total exempt count 30 → 31. The alert's record read is counted separately, as a read failure.
  • RetentionReevaluationLiveTests' hand-arm step writes both darling_armed and scheduled, which is how a real re-hold is exercised. A DBA's bare scheduled => true is a converge (logged), not a re-hold.
  • For #4186 gate follow-ups: release needs a second start after an outage, plus three smaller holds #4300's first item: the repair the Periodic pass launches runs the same walk, seam included.
  • Release-note facts:
    • on a downgrade, the previous version behaves as it does today;
    • the first start after the upgrade still runs the old armed job once, before the service converges it.

Findings during development

  • The third raw job: RawTierCoverage in TimescaleSupport.cs names exactly three raw relations — query_stats, procedure_stats, and query_store_stats.
  • Does lock_timeout apply inside run_job? Yes. With a second session holding ACCESS EXCLUSIVE on the oldest chunk for 30s, CALL run_job($1) wrapped in BEGIN; SET LOCAL lock_timeout='5s'; ...; COMMIT; raised ERROR: canceling statement due to lock timeout at 5.07s elapsed, not after the full 30s hold. Chunk count was unchanged after the failed call.
  • Can a crash leave a retry behind? Yes, safely, in both a clean backend termination and a SIGKILL crash-restart. Blocking CALL run_job on a held lock, then ending the blocked backend (pg_terminate_backend in one trial, SIGKILL in another, with a confirmed crash-restart log line in the SIGKILL trial), left job_stats.total_runs/last_run_started_at unset and the chunk count unchanged over a 10-minute, 20-poll window in both trials. The next hourly pass is the retry path; nothing needed manual cleanup.
  • The purge primitive's original SQL failed on every call. RunRetentionPurgeJobSql sent BEGIN; SET LOCAL lock_timeout = '5s'; CALL run_job($1::integer); COMMIT; as one NpgsqlCommand with a positional parameter. Npgsql's extended (prepared-statement) protocol rejects multiple statements in one command (SqlState 42601). The existing lock-timeout pin only asserted false + elapsed<20s, which the 42601 also satisfies, so nothing had shown the success path. Fixed by using an explicit NpgsqlTransaction with SET LOCAL lock_timeout and CALL run_job($1::integer) as separate commands, then CommitAsync. RunRetentionPurgeJobAsync now returns a RetentionPurgeOutcome(bool Ran, string? SqlState) instead of a bare bool, so a caller can see the SqlState when a run doesn't happen.

CHANGELOG entry

SECTION: Fixed
ENTRY:

Adds the SQL primitives for #4299's (d') design: the three raw
retention jobs (query_stats, procedure_stats, query_store_stats) move
their armed/held verdict off TimescaleDB's own scheduled flag onto
config->>'darling_armed', so the service can trigger the purge itself
instead of TimescaleDB's own job runner racing the service at start.

- ConvergeRawArmedStateSql: forces scheduled = false unconditionally
  and merges (||, never a replacing =) darling_armed into config.
- RawArmedReadExpression / RawArmedStateSql: the shared read that
  treats a missing darling_armed key as held, never armed.
- Five unit pins in RawArmedStateConvergeTests.cs.

Scope: lane 4299-1a only (converge SQL + read helper + unit pins).
Does not wire these into EnsureRetentionPoliciesAsync's sweep loop,
add the run_job primitive, or touch existing scheduled-based tests --
that is lane 4299-1b, next on this branch.
…e fresh (#4299 L4)

Confirmed #4186's seam-window fix (MaterializationHoleScanWindows) already
extends the scan's LOWER bound for a stitched successor's un-materialized
seam tail, but that is scoped to the successor/legacy seam case, not the
general raw-sourced case. The remaining gap: the ORDINARY window's lower
bound for a raw-sourced target (query_stats, procedure_stats,
query_store_stats) still clamped to utcNow minus MaterializationHoleScanSpanFor,
a horizon that assumed raw purged on its own schedule. Under #4299's
variant (d prime) the three raw jobs never run on a schedule at all, so raw
can legitimately hold rows far older than that clamp while the service-
triggered purge has not fired yet, and the old clamp would leave a real
hole below it unscanned indefinitely.

Fix: IsRawSourced(source) picks out the three RawTierCoverage relations;
for those, RawFloorHorizonAsync reads the raw table's own min(time column)
as the scan's lower bound instead of the time-based horizon (falling back
to the old horizon when raw is empty, where it is harmless).

Also builds the function L2's trigger will call as its gate:
HoleFreeThroughAsync(connection, target, materialization, dropFrom, dropTo,
deferredRanges) -- a FRESH scan of the exact range a purge would drop, never
a reuse of a stale repair result, that also treats any deferred range
(from the same pass's CapMaterializationHoleRepairs) overlapping the drop
window as a hole even though a bare scan of already-repaired buckets would
read clean.

Pins added to MaterializationHoleRepairTests.cs:
- IsRawSourced_IsTrueOnlyForTheThreeRawTierCoverageRelations
- DeferredRangeOverlappingDropWindow_WouldBlockTheGate_ByHalfOpenIntervalOverlap
Both RED on dev (the members don't exist there -- confirmed by grep).

No migrations. Only TimescaleSupport.MaterializationHoles.cs and its test
file touched, per brief -- did not touch TimescaleSupport.cs (4299-1a's file).
@erikdarlingdata erikdarlingdata changed the title DO NOT MERGE: raw retention jobs converge onto a config-key armed verdict (#4299) Service-triggered raw purge gate (#4299) Sep 26, 2026
For the three raw retention jobs (query_stats, procedure_stats,
query_store_stats — RawTierCoverage), EnsureRetentionPoliciesAsync now
runs ConvergeRawArmedStateSql instead of ArmRetentionPolicySql /
HoldRetentionPolicySql: scheduled stays permanently false and the
armed/held verdict is written to config->>'darling_armed' instead.
Every other retention job keeps today's scheduled semantics unchanged.

Because this runs inside the sweep (called at Startup and from the
hourly Periodic pass per #3812), the converge now happens on every
pass — a DBA's alter_job(scheduled => true) on a raw job gets reverted
to scheduled = false with the armed key restored within the hour, or
at the next start.

The prior-state read (used to tell a transition from a re-assertion,
for the Armed/ReHeld log lines and tally) now reads RawArmedStateSql
for a raw relation instead of RetentionPolicyScheduledSql, so the
transition detection stays correct under the new verdict.

Adds RawRelations (RawTierCoverage's relation names) as the lookup the
sweep uses to branch.

Not in this commit: CALL run_job — the raw jobs never purge yet.
…duled-based pins off query_stats (#4299)

The indeterminate-coverage branch previously left everything alone for
a raw relation too. Per the ruling (the hourly converge runs every
pass, unconditionally), a raw relation's scheduled flag now still
converges to false even when coverage could not be measured this pass
- only the darling_armed key is left untouched on that path.

Test migration (query_stats is now one of the three never-scheduled
raw relations, so it can no longer stand in for a scheduled-semantics
pin):
- TimescaleContinuousAggregateTests: ArmRetentionPolicy_Targets... and
  HoldRetentionPolicy_IsTheMirrorOfArming... now render against
  QueryStatsHourlyView (a non-raw relation) instead of query_stats.
  Assertions kept, not deleted.
- RetentionReevaluationTests (RetentionReevaluationLiveTests): added
  RawArmedAsync (reads config->>'darling_armed' via RawArmedStateSql)
  alongside the existing ScheduledAsync, and every held/armed
  assertion on HeldRelation (query_stats) now checks the armed key;
  scheduled is separately asserted to stay false throughout, proving
  the raw relation never reaches TimescaleDB's own scheduler.
- TheArmAndHoldStatements_StayUnconditional...'s source-order anchor
  updated for the new ternary read (isRawRelation ? RawArmedStateSql :
  RetentionPolicyScheduledSql).
Four live xUnit pins in a new RawRetentionServiceTriggeredLiveTests class,
proving:
1. A product-created raw retention job converges from the old
   scheduled=true/no-config shape onto scheduled=false with
   config->>'darling_armed' recorded, and drop_after survives.
2. A DBA's alter_job(scheduled => true) against an already-converged raw
   job is reverted by the very next sweep, regardless of the coverage
   verdict that pass reaches.
3. A non-raw retention relation keeps today's arm/hold semantics
   (scheduled flag, no darling_armed key).
4. A Covered raw job reads darling_armed = 'true' but scheduled stays
   false (never true, which is how the pre-fix code armed a raw
   relation).

All four verified GREEN in-process (dotnet Darling.Tests.dll, Release,
WindowsDesktop entry stripped) against a real TimescaleDB 2.30.1/PG18
rig, and RED against 8516650 (has the SQL, no wiring).
…s scheduled (#4299)

query_store_stats and query_stats are raw relations under #4299's design: their
armed verdict lives in config->>'darling_armed', converged through
ConvergeRawArmedStateSql/RawArmedStateSql, while their 'scheduled' column is
unconditionally driven to false. The test's IsArmedAsync still read only
'scheduled', so it always saw false for these two relations regardless of the
real armed state. Mirrors the product's own branch (RawRelations.Contains) so
the test and the product read the same column for the same relation. No
assertion intent changed.
…arling_armed (#4299)

RollupCreatedOverExistingHistory_ReadsEmptyBelowItsFloor_UntilTheBackfillFixesIt
and InterruptedBackfill_DoesNotLookComplete_AndTheArmingGateRefusesUntilItGenuinelyIs
both asserted query_stats' retention job's 'scheduled' column via a literal count
query (ArmedRawPolicySql). query_stats is a raw relation under #4299's design:
its armed verdict lives in config->>'darling_armed' via RawArmedStateSql, while
'scheduled' is unconditionally driven to false. Replaced ArmedRawPolicySql's two
call sites with a RawArmedAsync helper that reads through RawArmedStateSql. Kept
ArmedHourlyPolicySql/its call site unchanged: query_stats_interval_hourly is not
a raw relation and still uses scheduled semantics. No assertion intent changed.
…#4299)

Pass 1b's hand-arm (the runbook-forbidden alter_job the test uses to set up
a re-hold) only wrote scheduled => true. Under #4299 (d') query_stats' real
armed state lives in config->>'darling_armed', read via RawArmedStateSql;
EnsureRetentionPoliciesAsync's prior-state read for a raw relation consults
that key, not scheduled. The hand-arm therefore never registered as armed
from the sweep's point of view, so pass 1b's transition to re-held never
happened (re-held stayed 0). Fixed the hand-arm to also merge
darling_armed=true into the job's config, matching what
ConvergeRawArmedStateSql would have written, and added an assertion that it
took. Assertion intent unchanged; this is completing the raw-relation
migration this test file had already started (RawArmedAsync existed but the
setup step feeding it a transition to detect did not).
…heduled (#4299)

query_stats and procedure_stats are raw relations under #4299's design: their
armed verdict lives in config->>'darling_armed' via RawArmedStateSql, while
'scheduled' is unconditionally driven to false. IsRetentionScheduledAsync only
read 'scheduled', so it always saw false for both raw relations in
RawPurge_ArmsOffSuccessorHourlyCoverage_NotTheEmptyLegacyOnes regardless of the
real armed state. Now branches on TimescaleSupport.RawRelations, mirroring the
product's own branch, so the test reads the same column the product reads for
the same relation. No assertion intent changed.
…d raw purge (#4299)

Lane 4299-1c on the shared branch.

- RunRetentionPurgeJobSql/RunRetentionPurgeJobAsync: the service's own
  CALL run_job(id) trigger for a raw retention job under variant (d'),
  wrapped in BEGIN; SET LOCAL lock_timeout; CALL run_job($1); COMMIT;
  with a bounded lock_timeout (RunRetentionPurgeJobLockTimeout, 5s) and a
  named CommandTimeout (RunRetentionPurgeJobTimeoutSeconds =
  SetupTimeoutSeconds). A timeout or any other failure is logged at
  Warning, best-effort ROLLBACK recovers the connection, and the method
  returns false so the caller's next hourly pass retries. Not called
  from anywhere yet (L2's job).
- Two rig probes, run on real TimescaleDB 2.30.1/PostgreSQL 18
  containers, both PASS:
  1. lock_timeout applies INSIDE run_job: a second session held ACCESS
     EXCLUSIVE on the target chunk, CALL run_job wrapped in
     BEGIN/SET LOCAL/COMMIT raised 55P03 lock_not_available at ~5s (not
     30s of blocking), chunk count unchanged after. Also proven as a
     new xUnit live pin, RawRetentionPurgeJobLiveTests, GREEN in-process
     on a real rig (0.7s).
  2. crash safety: started run_job on a never-scheduled job holding a
     lock, then ended it two ways in parallel containers -
     pg_terminate_backend and SIGKILL (forces a postmaster
     crash-restart). After each: 10 minutes of 30s polling showed no
     scheduler run of the job (total_runs/last_run_started_at
     unchanged) and the chunk count held at 40 in both trials.
- Docs: RetentionArmSafetySql's XML doc and the "not load-bearing"
  comment in TimescaleSupport.MaterializationHoles.cs (plus its twin in
  DarlingWorker.cs) now state that the raw purge only runs when the
  service triggers it after a Covered verdict and a finished repair
  under the current pg_postmaster_start_time(), naming both exceptions:
  the first start after the upgrade, and a DBA's alter_job/run_job
  before the next converge.

Build: Darling.Tests and Lite.Tests both 0/0 warnings/errors with
-p:EnableWindowsTargeting=true. StorageCommandTimeoutTests,
LiveCleanupConversionRatchetTests, StartupCommandTimeoutTests,
DocCommentHygiene, AlertReadFailureSurfaceTests all green in-process.
TriggerRawPurgeCoreAsync(connection, logger, cancellationToken) takes the
exact body of the private instance method, unchanged, so Darling.Tests
(already reachable via this project's InternalsVisibleTo) can drive the
trigger directly against a live store for the end-to-end pins. The
instance method is now a one-line forward. No behaviour change.
…, surface SqlState

RunRetentionPurgeJobSql sent BEGIN; SET LOCAL lock_timeout; CALL run_job($1::integer); COMMIT;
as one NpgsqlCommand with a positional parameter. Npgsql's extended protocol rejects that with
42601 ("cannot insert multiple commands into a prepared statement"), so the CALL never ran and
RunRetentionPurgeJobAsync always returned false, silently. The existing lock-timeout pin only
asserted false + elapsed<20s, so it passed on the 42601 without ever exercising the real path.

Proven on the rig with psql (unblocked): BEGIN; SET LOCAL lock_timeout='5s'; CALL run_job(<id>);
COMMIT; succeeds and drops the chunk. Fixed by opening an explicit NpgsqlTransaction and sending
SET LOCAL and CALL as separate commands on it.

RunRetentionPurgeJobAsync now returns a RetentionPurgeOutcome(bool Ran, string? SqlState) instead
of a bare bool, so a caught failure states which SQLSTATE it was. TriggerRawPurgeCoreAsync (the
one caller) logs it.

Pins (live, Darling/Darling.Tests/RawRetentionPurgeJobLiveTests.cs):
- RunRetentionPurgeJob_Unblocked_RunsAndDropsTheOldChunk (new): unblocked with a chunk older than
  drop_after, asserts Ran==true and the chunk count drops. RED on 0c3ed6a: false, count unchanged.
- RunRetentionPurgeJob_BlockedOnAConflictingLock_...: now also asserts SqlState=="55P03", so it
  can never pass on a 42601 again. RED on 0c3ed6a: SqlState was 42601, not surfaced by the old
  bare-bool return.

Both pins run GREEN in-process on this branch (Darling.Tests.dll, macOS, TimescaleDB 2.30.1).
Copies the lane-4299-2d draft RawPurgeTriggerLiveTests.cs into the repo.
All five [Fact]s pass in-process against the fixed RunRetentionPurgeJobAsync
from 9329306: no product or seed changes were needed.
…tency

Extracts DarlingWorker.ShouldLaunchMaterializationHoleRepair as a pure
internal static method (behaviour unchanged) so a second service's
relaunch decision can be asserted directly. Adds
RawPurgeTriggerLiveTests.TwoServices_RelaunchGuardHolds_PurgeRunsOnceOnly:
two independent NpgsqlDataSources against the same scratch store prove
Service B sees Service A's epoch stamp, decides not to launch a repair,
re-stamps 0 rows, and that TriggerRawPurgeCoreAsync called from both
services in turn drops chunks once and is a no-op the second time.
…r raw jobs

- RetentionHoldReadSql now projects config->>'darling_armed' as an extra
  column; ReadRetentionHoldReadingsAsync picks darling_armed for one of
  the three raw relations (TimescaleSupport.RawRelations) and keeps
  j.scheduled for every other retention policy, so the Retention Held
  alert judges a raw job by its real verdict instead of its permanent
  scheduled=false.
- JobCadenceReadSql excludes the three raw retention jobs from the #2136
  cadence reading (the minimum-safe exclusion; cadence-from-trigger is a
  follow-up lane's work).
- ApplyPolicyJobsStuckAsync's not-scheduled WARNING now says a raw job is
  unscheduled by design (#4299), not a #1680/#1877 coverage hold, when the
  hypertable is one of the three raw relations.
…text branch

- RetentionHoldAndCadenceRawRedirectLiveTests (live, own-store): darling_armed
  true/false on RetentionHoldReadSql, and JobCadenceReadSql's absence of a raw
  job even after a successful triggered run.
- DarlingSelfAlertTests: the M4 pin (a raw job flagged stuck is never re-armed,
  GREEN on both commits) and the WARNING-text pin (unscheduled-by-design for a
  raw relation, RED before this lane's branch).
@erikdarlingdata erikdarlingdata changed the title Service-triggered raw purge gate (#4299) Raw retention purges run only when the service triggers them, never at PostgreSQL start (#4299) Sep 26, 2026
…solvable successor (#4391)

The service-triggered raw purge (#4299) read the coverage sweep's
standing darling_armed flag to decide whether to drop chunks. That
flag is left unchanged whenever a probe cannot measure coverage this
pass, so a purge could run on a stale Covered verdict from an earlier
hour even though this pass's own coverage read came back Unknown.
Separately, when a successor's continuous aggregate could not be
resolved (mid-rebuild, renamed, or dropped), the hole check skipped
it and treated the raw relation as hole-free.

The trigger now measures coverage fresh in the same pass via
IsRawTierDropSafeAsync, which answers true only on a Covered verdict
from this call; darling_armed keeps its existing meaning for readers
and alerts and is untouched. An unresolvable successor now blocks
the purge and records the outcome as gate_unknown instead of being
skipped.

Added two live pins to RawPurgeTriggerLiveTests: a stale
darling_armed=true with coverage measured Unknown this pass, and a
Covered/current-epoch state whose hole-check successor cannot be
resolved. Both drop no chunks on the fixed code.
…nreadable one as never-written (#4299, #4391)

- TimescaleSupport.cs: add ReadRawLastPurgeStateAsync, returning a tri-state
  RawLastPurgeReadState (Present/NeverWritten/ReadFailed) alongside the record.
  ReadRawLastPurgeOutcomeAsync becomes a thin wrapper so its seven existing
  callers keep compiling unchanged. RawPurgeOverHorizonReading gains a
  trailing LastPurgeReadFailed flag (default false, so every existing
  constructor call still compiles).
- DarlingSelfAlertEvaluator.cs: the Raw Purge Over Horizon alert now also
  fires when the last-purge record is more than two hours old (twice the
  hourly tick the purge trigger rides on), even if its outcome was "ran" --
  an old ran record must not keep the alert quiet forever once the trigger
  itself has stopped running. A read failure is counted (via the existing
  swallowed-read counter) and skipped: it neither fires nor resolves the
  alert, leaving the standing state exactly where it was. Reason text gains
  gate_unknown and gate_error arms.
- DarlingSelfAlertTests.cs: pins for the stale-record fire (with its detail
  text), the recent-record non-fire, a read failure while the alert is
  active (must not resolve, must count), a read failure while inactive
  (must not fire), and the two new reason-text arms.
…he raw purge trigger's own gate failure (#4299, #4391)

- RepairEpochStampAllowed(summary) allows the stamp only when the materialization-hole repair pass had zero isolated failures; deferred/remaining holes do not block it (the trigger's own fresh hole scan covers those).
- TriggerRawPurgeCoreAsync's per-relation catch now logs at Warning and records the outcome as gate_error (with SqlState when available) instead of silently logging at Information with nothing recorded.
- Adds RawPurgeTriggerGateErrorTests.cs: a pure pin for RepairEpochStampAllowed and a live pin that forces the trigger's hole scan to fail with a lock timeout and asserts no chunk drops and the outcome records gate_error/57014.
…riodic repair (#4299, #4391)

- ConvergeRawArmedStateSql now writes only when the job is scheduled or its darling_armed key differs from the verdict, matching RawRepairEpochStampSql's guard style; the Unknown branch's schedule-only revert is guarded the same way, and logs one Information line when a DBA's alter_job(scheduled => true) is reverted.
- The Periodic relaunch now reads all three raw jobs' repair epoch and relaunches when ANY is stale, not just the first; the log line names the stale relations.
- The Periodic pass's own repair launch is kept in a field and awaited in the same shutdown drain as the start-path launch, so it is neither lost nor unobserved on shutdown.
- RecordRawLastPurgeOutcomeAsync's record is now built server-side with jsonb_build_object and explicit casts, never string-interpolated.
…urge read (#4299, #4391)

RawPurgeConvergeLiveTests.cs: scheduled=false holds after Covered/Short/Unknown when a DBA re-arms a raw job by hand; a second converge pass on an already-converged store writes nothing and logs the revert line exactly once, on the pass that actually reverted something; the tri-state last-purge read distinguishes NeverWritten, ReadFailed (a malformed value) and Present (a real trigger pass) live.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant