Repository navigation
Stop the compression detector reading TimescaleDB's never-ran sentinel as year 1 (#1760) - #1781
Merged
Merged
Conversation
…l as year 1 (#1760) The nightly flake was not a slow runner. Two measured facts: 1. timescaledb_information.job_stats.last_run_started_at is -infinity, NOT NULL, for a job that has never run. Npgsql maps that to DateTime.MinValue, so IsCompressionJobStuck computed a ~739,000-day elapsed that clears every StuckRunningBound and flagged a perfectly healthy job as stuck. 2. That was reachable because job_status and last_run_started_at come from INDEPENDENT sources in TimescaleDB's own view: job_status is CASE WHEN pg_stat_activity.state = 'active' THEN 'Running', joined on application_name, while last_run_started_at is bgw_job_stat.last_start. A job's FIRST run reads Running while its start time is still the sentinel, so the window is structural rather than hypothetical. This is a production defect, not only a test one: the self-heal would re-arm a job that was running fine. StuckCompressionJobsSql becomes a const (the RearmJobSql idiom) and NULLIFs the sentinel, so both -infinity tests run in SQL rather than through Npgsql's infinity mapping. IsCompressionJobStuck rejects DateTime.MinValue as a second line of defence. The settle-wait polled next_start only - one of the two arms the detector evaluates - so "settled" and "the assertion will pass" were different statements. Worse, the value it polled was the one the caller's own alter_job had just written: instrumented, the original wait returned after POLLS=1 DELAYS=0. It waited for nothing, which is why the earlier hardening did not help and why no timeout increase could have. It now polls ReadStuckCompressionJobsAsync itself, so there is one predicate rather than two copies that drift, and it reports the detector's own reason on timeout. Leg (1) also gains the assertion it always claimed to make: the detector is failure-isolated, so a broken query returns an EMPTY list and DoesNotContain passes just as happily against SQL that never compiled. The test now runs the production const directly, where an error throws, and requires the job to be present. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ders (#1781) CI went red: 42883: function by_range(unknown) does not exist. The new test hand-rolled by_range('collection_time'), whose one-argument form TimescaleDB 2.28 (the local rig) accepts and the older version CI's fixture carries does not. Use CreateHypertableSql / EnableCompressionSql / AddCompressionPolicySql instead. Those are the forms the rest of this class already exercises green on CI, and the test is about the catalog's behaviour rather than a second dialect of the same DDL. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…on-settle-predicate # Conflicts: # CHANGELOG.md
…on-settle-predicate # Conflicts: # CHANGELOG.md
erikdarlingdata
enabled auto-merge
July 28, 2026 04:21
CI red: 42883: function by_range(unknown, interval) does not exist. Switching to the product's SQL builders last round was necessary but not sufficient - the two-argument form failed too, which pointed at resolution rather than syntax. by_range is in the PUBLIC schema (verified: pg_proc join pg_namespace -> nspname 'public'). A session whose search_path omits public therefore cannot resolve it, create_hypertable never resolves its argument, and the whole statement dies naming the inner function. This test was the only gated one in the class that did not call MigrateAsync first, which is what every sibling relies on to establish the path. Reproduced locally rather than guessed: with the connection pinned to SearchPath=collect,config - CI's shape - the test fails with CI's exact error, and passes with the migrate call restored. My earlier rig passed only because its connection carried the default "$user", public. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…on-settle-predicate # Conflicts: # CHANGELOG.md
…on-settle-predicate # Conflicts: # CHANGELOG.md
erikdarlingdata
added a commit
that referenced
this pull request
Jul 28, 2026
Resolved to get CI to run at all, not as an arm-time step: a CONFLICTING PR has no merge ref, and pull_request workflows build the merge ref -- so Build and Claude Auto Review could not start on this PR while the conflict stood, while other branches' runs kept firing normally. That is what made this look like an Actions outage for ~25 minutes. CHANGELOG.md was the only conflict, and only in the link-ref block. Resolved KEEP-BOTH with no re-sorting: this branch's #1665/#1788 refs and dev's #1781/#1783/#1786/#1792 refs all retained, in the order they appeared. No entry text on either side was touched. Everything else auto-merged. Verified rather than assumed, because the dangerous case here is textual cleanliness hiding a semantic break: full solution rebuild 0 warnings / 0 errors, Darling suite 3549 passed / 0 failed. The one real collision this class of merge produced -- #1776's new live-store hygiene guard versus this branch's new own-store live test -- was found the same way and fixed in 59f877a before this merge. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
erikdarlingdata
added a commit
that referenced
this pull request
Jul 28, 2026
Conflicts were a using-directive collision in TimescaleSupportTests.cs and the routine CHANGELOG stacking; both resolved keep-both, dev's ordering untouched. Test union verified exact: 27 = dev's 19 + mine, nothing dropped. TimescaleSupport.cs auto-merged CLEANLY and that was the actual hazard. #1760 (on dev via #1781) established that TimescaleDB stores -infinity, not NULL, as the never-ran sentinel in last_run_started_at, that Npgsql maps it to DateTime.MinValue, and that job_status comes from pg_stat_activity INDEPENDENTLY -- so a policy's first run reads Running while its start is still the sentinel. The compression-activity query #1778 added in parallel reads the same column and had no such guard, so the merged tree would have logged a healthy first run as having gone for ~739,000 days. Neither branch could catch it alone; it lives only in the seam. Applies the same NULLIF to CompressionActivitySql and the same MinValue guard to CompressionActivity.RunningFor. The existing zero-clamp there does not cover it -- the sentinel yields a huge POSITIVE elapsed.
This was referenced Jul 28, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #1760.
The flake was a production defect, not a slow runner.
ReadStuckCompressionJobsAsyncis the compression self-heal's detector (#1585/#1586), and it flagged healthy jobs as stuck for the whole of their first run.Root cause, measured
Both facts below were measured against a real PostgreSQL + TimescaleDB 2.28.1 cluster, not inferred.
1.
last_run_started_atis-infinity, not NULL, for a job that has never run. Freshly added policy has nobgw_job_statrow at all (every column NULL); it isalter_jobthat materialises the row carrying the sentinel:Npgsql maps
-infinitytoDateTime.MinValue, soIsCompressionJobStuckcomputednowUtc - DateTime.MinValue— roughly 739,000 days — which clears everyStuckRunningBound(floor 2h). The job was reported "stuck in the Running state for 388,000,000 minutes".2. Why that was reachable at all.
job_statusandlast_run_started_atcome from independent sources in TimescaleDB's own view definition:job_statusis derived frompg_stat_activity;last_run_started_atisbgw_job_stat.last_start. A job's FIRST run therefore readsRunningwhile its start time is still the sentinel. The window is structural, not hypothetical — and every job on a fresh CI cluster is in exactly that state.ApplyCompressionPolicyAsynccreates a policy for all 26 collector tables at once and the test seeds a 40-day-oldwait_statsrow, so several jobs are starting real compression work around the moment the detector reads. That is the interleaving the nightly hit.Why the previous hardening could not have worked
The issue asked to distinguish two candidates. Both are settled by evidence rather than argument.
Candidate 2 (the wait times out and returns quietly) is refuted outright. The wait does not return quietly on timeout —
TimescaleSupportTests.cs:645asserts loudly with a distinctive "never settled to a real next_start" message. The observed nightly failure wasAssert.DoesNotContain() Failure, a different assertion entirely.Candidate 1 is confirmed, and worse than described. The wait polled
next_start <> '-infinity'— one of the two arms the detector evaluates — so "settled" and "the assertion will pass" were different statements. But the value it polled was the one the caller's ownalter_job(next_start => now() + interval '1 hour')had written one statement earlier. Instrumented (M3 below), the original wait returned afterPOLLS=1 DELAYS=0. It never waited at all, and the 30s deadline was never approached — which is precisely why no timeout increase could ever have helped.The fix
Production.
StuckCompressionJobsSqlbecomes apublic const(the existingRearmJobSqlidiom in the same file) andNULLIFs the sentinel, so both-infinitytests run in SQL rather than through Npgsql's infinity mapping — which is the discipline that comment already claimed fornext_start.IsCompressionJobStuckadditionally rejectsDateTime.MinValue, so a future caller reading the column un-guarded cannot resurrect the false positive.Test. The settle-wait now polls
ReadStuckCompressionJobsAsyncitself. That is predicate identity by construction — one definition, not two copies that drift — so settled-according-to-the-wait is settled-according-to-the-assertion. It reports the detector's ownReasonon timeout instead of assumingnext_start. No retry wrapper, no timeout increase.Leg (1) also gains the assertion it always claimed to make.
ReadStuckCompressionJobsAsyncis failure-isolated: a broken query is swallowed and returns an EMPTY list, soDoesNotContainalone passed just as happily against SQL that never compiled — the one thing that leg exists to prove. It now runs the production const directly, where a syntax or column error throws, and requires the job to be present in the result.Mutation table
Each row was applied, built, and watched red.
NULLIFfromStuckCompressionJobsSqlStuckCompressionJobsSql_GuardsTheNeverRanSentinel,StuckCompressionJobsSql_NeverRunJob_ReadsNullLastRunStartedAt_AgainstDevPostgresDateTime.MinValueguard fromIsCompressionJobStuckIsCompressionJobStuck_RunningWithNeverRanSentinel_IsNotStucknext_start-only wait, instrumentedPOLLS=1 DELAYS=0M1 is the load-bearing one: the live test went red, which proves the sentinel really is
-infinityin the catalog. Had it been NULL, removing theNULLIFwould have changed nothing and the test would have stayed green — so the guard is demonstrably not dead code.M3 is the flake's mechanism made measurable rather than argued.
Verification
Darling.Testsfull suite against a live PostgreSQL 18 + TimescaleDB 2.28.1 rig: 3651 passed, 0 failed, 7 skipped.TimescaleSupportTestsalone: 19/19, including the 3 new tests (2 pure, 1 gated live).dotnet build Darling/Darling.Tests/Darling.Tests.csproj -c Debug -t:Rebuild— 0 Warning(s), 0 Error(s).The new live test is deterministic rather than timing-dependent: arming the job an hour out both materialises the stat row carrying the sentinel and guarantees the scheduler cannot run the job out from under the assertion. It asserts the raw column really does carry
-infinity(so the guard cannot pass for the wrong reason) and that the production query hands back NULL.Field relevance
Worth flagging for the release: on any store where a compression job is observed during its first run, the self-heal was re-arming a healthy job via
alter_job(next_start => now()). That is not data loss and it is self-limiting — the sentinel clears once the first run finishes — but it means "compression job stuck" alerts on a newly provisioned or freshly upgraded store may have been false, and the recorded reason ("stuck in the Running state for 388,000,000 minutes") is the signature to look for in field logs.🤖 Generated with Claude Code