Skip to content

Stop the compression detector reading TimescaleDB's never-ran sentinel as year 1 (#1760) - #1781

Merged
erikdarlingdata merged 8 commits into
devfrom
feature/1760-compression-settle-predicate
Jul 28, 2026
Merged

erikdarlingdata merged 8 commits into
devfrom
feature/1760-compression-settle-predicate

Conversation

@erikdarlingdata

Copy link
Copy Markdown
Owner

Closes #1760.

The flake was a production defect, not a slow runner. ReadStuckCompressionJobsAsync is the compression self-heal's detector (#1585/#1586), and it flagged healthy jobs as stuck for the whole of their first run.

Root cause, measured

Both facts below were measured against a real PostgreSQL + TimescaleDB 2.28.1 cluster, not inferred.

1. last_run_started_at is -infinity, not NULL, for a job that has never run. Freshly added policy has no bgw_job_stat row at all (every column NULL); it is alter_job that materialises the row carrying the sentinel:

 job_id | next_start_neg_inf |          next_start          | job_status | last_run_started_at | lrs_neg_inf | lrs_is_null
--------+--------------------+------------------------------+------------+---------------------+-------------+-------------
   1000 | f                  | 2026-07-28 00:21:41.86261-04 | Scheduled  | -infinity           | t           | f

Npgsql maps -infinity to DateTime.MinValue, so IsCompressionJobStuck computed nowUtc - DateTime.MinValue — roughly 739,000 days — which clears every StuckRunningBound (floor 2h). The job was reported "stuck in the Running state for 388,000,000 minutes".

2. Why that was reachable at all. job_status and last_run_started_at come from independent sources in TimescaleDB's own view definition:

CASE WHEN pgs.state = 'active' THEN 'Running' ... END AS job_status
...
LEFT JOIN pg_stat_activity pgs ON pgs.datname = current_database()
                              AND pgs.application_name = j.application_name

job_status is derived from pg_stat_activity; last_run_started_at is bgw_job_stat.last_start. A job's FIRST run therefore reads Running while its start time is still the sentinel. The window is structural, not hypothetical — and every job on a fresh CI cluster is in exactly that state.

ApplyCompressionPolicyAsync creates a policy for all 26 collector tables at once and the test seeds a 40-day-old wait_stats row, so several jobs are starting real compression work around the moment the detector reads. That is the interleaving the nightly hit.

Why the previous hardening could not have worked

The issue asked to distinguish two candidates. Both are settled by evidence rather than argument.

Candidate 2 (the wait times out and returns quietly) is refuted outright. The wait does not return quietly on timeout — TimescaleSupportTests.cs:645 asserts loudly with a distinctive "never settled to a real next_start" message. The observed nightly failure was Assert.DoesNotContain() Failure, a different assertion entirely.

Candidate 1 is confirmed, and worse than described. The wait polled next_start <> '-infinity' — one of the two arms the detector evaluates — so "settled" and "the assertion will pass" were different statements. But the value it polled was the one the caller's own alter_job(next_start => now() + interval '1 hour') had written one statement earlier. Instrumented (M3 below), the original wait returned after POLLS=1 DELAYS=0. It never waited at all, and the 30s deadline was never approached — which is precisely why no timeout increase could ever have helped.

The fix

Production. StuckCompressionJobsSql becomes a public const (the existing RearmJobSql idiom in the same file) and NULLIFs the sentinel, so both -infinity tests run in SQL rather than through Npgsql's infinity mapping — which is the discipline that comment already claimed for next_start. IsCompressionJobStuck additionally rejects DateTime.MinValue, so a future caller reading the column un-guarded cannot resurrect the false positive.

Test. The settle-wait now polls ReadStuckCompressionJobsAsync itself. That is predicate identity by construction — one definition, not two copies that drift — so settled-according-to-the-wait is settled-according-to-the-assertion. It reports the detector's own Reason on timeout instead of assuming next_start. No retry wrapper, no timeout increase.

Leg (1) also gains the assertion it always claimed to make. ReadStuckCompressionJobsAsync is failure-isolated: a broken query is swallowed and returns an EMPTY list, so DoesNotContain alone passed just as happily against SQL that never compiled — the one thing that leg exists to prove. It now runs the production const directly, where a syntax or column error throws, and requires the job to be present in the result.

Mutation table

Each row was applied, built, and watched red.

# Mutation Result Test that went red
M1 Remove NULLIF from StuckCompressionJobsSql RED (2) StuckCompressionJobsSql_GuardsTheNeverRanSentinel, StuckCompressionJobsSql_NeverRunJob_ReadsNullLastRunStartedAt_AgainstDevPostgres
M2 Remove the DateTime.MinValue guard from IsCompressionJobStuck RED (1) IsCompressionJobStuck_RunningWithNeverRanSentinel_IsNotStuck
M3 Restore the original next_start-only wait, instrumented RED, reporting POLLS=1 DELAYS=0 measurement probe, removed before commit

M1 is the load-bearing one: the live test went red, which proves the sentinel really is -infinity in the catalog. Had it been NULL, removing the NULLIF would have changed nothing and the test would have stayed green — so the guard is demonstrably not dead code.

M3 is the flake's mechanism made measurable rather than argued.

Verification

  • Darling.Tests full suite against a live PostgreSQL 18 + TimescaleDB 2.28.1 rig: 3651 passed, 0 failed, 7 skipped.
  • TimescaleSupportTests alone: 19/19, including the 3 new tests (2 pure, 1 gated live).
  • Build: dotnet build Darling/Darling.Tests/Darling.Tests.csproj -c Debug -t:Rebuild — 0 Warning(s), 0 Error(s).

The new live test is deterministic rather than timing-dependent: arming the job an hour out both materialises the stat row carrying the sentinel and guarantees the scheduler cannot run the job out from under the assertion. It asserts the raw column really does carry -infinity (so the guard cannot pass for the wrong reason) and that the production query hands back NULL.

Field relevance

Worth flagging for the release: on any store where a compression job is observed during its first run, the self-heal was re-arming a healthy job via alter_job(next_start => now()). That is not data loss and it is self-limiting — the sentinel clears once the first run finishes — but it means "compression job stuck" alerts on a newly provisioned or freshly upgraded store may have been false, and the recorded reason ("stuck in the Running state for 388,000,000 minutes") is the signature to look for in field logs.

🤖 Generated with Claude Code

erikdarlingdata and others added 3 commits July 27, 2026 23:34
…l as year 1 (#1760)

The nightly flake was not a slow runner. Two measured facts:

1. timescaledb_information.job_stats.last_run_started_at is -infinity, NOT
   NULL, for a job that has never run. Npgsql maps that to DateTime.MinValue,
   so IsCompressionJobStuck computed a ~739,000-day elapsed that clears every
   StuckRunningBound and flagged a perfectly healthy job as stuck.

2. That was reachable because job_status and last_run_started_at come from
   INDEPENDENT sources in TimescaleDB's own view: job_status is
   CASE WHEN pg_stat_activity.state = 'active' THEN 'Running', joined on
   application_name, while last_run_started_at is bgw_job_stat.last_start.
   A job's FIRST run reads Running while its start time is still the
   sentinel, so the window is structural rather than hypothetical.

This is a production defect, not only a test one: the self-heal would re-arm
a job that was running fine.

StuckCompressionJobsSql becomes a const (the RearmJobSql idiom) and NULLIFs
the sentinel, so both -infinity tests run in SQL rather than through Npgsql's
infinity mapping. IsCompressionJobStuck rejects DateTime.MinValue as a second
line of defence.

The settle-wait polled next_start only - one of the two arms the detector
evaluates - so "settled" and "the assertion will pass" were different
statements. Worse, the value it polled was the one the caller's own alter_job
had just written: instrumented, the original wait returned after POLLS=1
DELAYS=0. It waited for nothing, which is why the earlier hardening did not
help and why no timeout increase could have. It now polls
ReadStuckCompressionJobsAsync itself, so there is one predicate rather than
two copies that drift, and it reports the detector's own reason on timeout.

Leg (1) also gains the assertion it always claimed to make: the detector is
failure-isolated, so a broken query returns an EMPTY list and DoesNotContain
passes just as happily against SQL that never compiled. The test now runs the
production const directly, where an error throws, and requires the job to be
present.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ders (#1781)

CI went red: 42883: function by_range(unknown) does not exist. The new test
hand-rolled by_range('collection_time'), whose one-argument form TimescaleDB
2.28 (the local rig) accepts and the older version CI's fixture carries does
not.

Use CreateHypertableSql / EnableCompressionSql / AddCompressionPolicySql
instead. Those are the forms the rest of this class already exercises green on
CI, and the test is about the catalog's behaviour rather than a second dialect
of the same DDL.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…on-settle-predicate

# Conflicts:
#	CHANGELOG.md
…on-settle-predicate

# Conflicts:
#	CHANGELOG.md
erikdarlingdata and others added 3 commits July 28, 2026 00:39
CI red: 42883: function by_range(unknown, interval) does not exist. Switching
to the product's SQL builders last round was necessary but not sufficient -
the two-argument form failed too, which pointed at resolution rather than
syntax.

by_range is in the PUBLIC schema (verified: pg_proc join pg_namespace ->
nspname 'public'). A session whose search_path omits public therefore cannot
resolve it, create_hypertable never resolves its argument, and the whole
statement dies naming the inner function. This test was the only gated one in
the class that did not call MigrateAsync first, which is what every sibling
relies on to establish the path.

Reproduced locally rather than guessed: with the connection pinned to
SearchPath=collect,config - CI's shape - the test fails with CI's exact error,
and passes with the migrate call restored. My earlier rig passed only because
its connection carried the default "$user", public.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…on-settle-predicate

# Conflicts:
#	CHANGELOG.md
…on-settle-predicate

# Conflicts:
#	CHANGELOG.md
@erikdarlingdata
erikdarlingdata merged commit b357510 into dev Jul 28, 2026
4 checks passed
erikdarlingdata added a commit that referenced this pull request Jul 28, 2026
Resolved to get CI to run at all, not as an arm-time step: a CONFLICTING PR
has no merge ref, and pull_request workflows build the merge ref -- so Build
and Claude Auto Review could not start on this PR while the conflict stood,
while other branches' runs kept firing normally. That is what made this look
like an Actions outage for ~25 minutes.

CHANGELOG.md was the only conflict, and only in the link-ref block. Resolved
KEEP-BOTH with no re-sorting: this branch's #1665/#1788 refs and dev's
#1781/#1783/#1786/#1792 refs all retained, in the order they appeared. No
entry text on either side was touched.

Everything else auto-merged. Verified rather than assumed, because the
dangerous case here is textual cleanliness hiding a semantic break: full
solution rebuild 0 warnings / 0 errors, Darling suite 3549 passed / 0 failed.
The one real collision this class of merge produced -- #1776's new live-store
hygiene guard versus this branch's new own-store live test -- was found the
same way and fixed in 59f877a before this merge.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@erikdarlingdata
erikdarlingdata deleted the feature/1760-compression-settle-predicate branch July 28, 2026 05:05
erikdarlingdata added a commit that referenced this pull request Jul 28, 2026
Conflicts were a using-directive collision in TimescaleSupportTests.cs and
the routine CHANGELOG stacking; both resolved keep-both, dev's ordering
untouched. Test union verified exact: 27 = dev's 19 + mine, nothing dropped.

TimescaleSupport.cs auto-merged CLEANLY and that was the actual hazard.
#1760 (on dev via #1781) established that TimescaleDB stores -infinity, not
NULL, as the never-ran sentinel in last_run_started_at, that Npgsql maps it
to DateTime.MinValue, and that job_status comes from pg_stat_activity
INDEPENDENTLY -- so a policy's first run reads Running while its start is
still the sentinel. The compression-activity query #1778 added in parallel
reads the same column and had no such guard, so the merged tree would have
logged a healthy first run as having gone for ~739,000 days. Neither branch
could catch it alone; it lives only in the seam.

Applies the same NULLIF to CompressionActivitySql and the same MinValue
guard to CompressionActivity.RunningFor. The existing zero-clamp there does
not cover it -- the sentinel yields a huge POSITIVE elapsed.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant