The observation
TimescaleSupportTests.EnsureRetentionPolicies_ConvergesAnOldHorizon_PreservingScheduledStateAndNextStart_AgainstDevPostgres failed once on a nightly with expected 17, got 0, then passed on re-run — on the same commit that had passed the same job ~30 minutes earlier on PR #2812.
It was called a flake. I could not determine the root cause, and I am not going to assert one. What follows is what I ruled out with evidence, and the reason the next occurrence will be just as undiagnosable unless one thing changes.
Ruled out
- Static initialisation order.
RetentionPolicies is built from RawTierCoverage and BaselineAggregates. Both are declared in the same class at lines 891 and 1912, before RetentionPolicies at 1949, so textual init order is correct and the collection cannot be empty. (An empty collection would have produced 0 with no error noise, which fit the symptom well — but it is not possible here.)
- Within-assembly parallelism.
TimescaleSupportTests carries [Collection("live-postgres")], so it is serialised against the other live classes.
- Cross-job database sharing. Both
build.yml and nightly.yml set DARLING_TEST_PG to Host=127.0.0.1;Port=5541, a per-runner instance. Two jobs cannot be stepping on one database.
What got 0 actually means
EnsureRetentionPoliciesAsync increments applied only after each relation's transaction commits. A return of 0 therefore means all 17 relations threw and every one landed in the per-relation catch — not that the loop was skipped.
Why it is undiagnosable, which is the real defect
That per-relation catch is:
catch (Exception ex) when (ex is not OperationCanceledException)
{
logger?.LogWarning("Retention policy for {Relation} ({DropAfter}) failed - ...: {Message}", relation, dropAfter, ex.Message);
}
and every call site in this test passes null for the logger:
Assert.Equal(RetentionPolicyCount, await TimescaleSupport.EnsureRetentionPoliciesAsync(connection, null, ct));
So all 17 exception messages are discarded by the ?. and the only surviving evidence is the bare count. There is no capturing ILogger anywhere in Darling.Tests to pass instead.
This is the same shape as #2801 and #2816: a failure that reports a plausible-looking value while destroying the one piece of information needed to act on it. A test that can only ever say "expected 17, got 0" cannot be triaged, so it will keep being labelled a flake whether or not it is one.
Suggested fix
Add a minimal capturing ILogger to Darling.Tests and pass it at the EnsureRetentionPoliciesAsync call sites in this test, then include the captured warnings in the assertion message. Then the next occurrence names the actual Postgres error on the first failure instead of costing another round.
Worth doing before concluding anything about flakiness, and worth prioritising because this sits on a code path under active change (#2809 retention work, #2811/#2816 fetch-phase work).
The observation
TimescaleSupportTests.EnsureRetentionPolicies_ConvergesAnOldHorizon_PreservingScheduledStateAndNextStart_AgainstDevPostgresfailed once on a nightly with expected 17, got 0, then passed on re-run — on the same commit that had passed the same job ~30 minutes earlier on PR #2812.It was called a flake. I could not determine the root cause, and I am not going to assert one. What follows is what I ruled out with evidence, and the reason the next occurrence will be just as undiagnosable unless one thing changes.
Ruled out
RetentionPoliciesis built fromRawTierCoverageandBaselineAggregates. Both are declared in the same class at lines 891 and 1912, beforeRetentionPoliciesat 1949, so textual init order is correct and the collection cannot be empty. (An empty collection would have produced0with no error noise, which fit the symptom well — but it is not possible here.)TimescaleSupportTestscarries[Collection("live-postgres")], so it is serialised against the other live classes.build.ymlandnightly.ymlsetDARLING_TEST_PGtoHost=127.0.0.1;Port=5541, a per-runner instance. Two jobs cannot be stepping on one database.What
got 0actually meansEnsureRetentionPoliciesAsyncincrementsappliedonly after each relation's transaction commits. A return of0therefore means all 17 relations threw and every one landed in the per-relation catch — not that the loop was skipped.Why it is undiagnosable, which is the real defect
That per-relation catch is:
and every call site in this test passes
nullfor the logger:So all 17 exception messages are discarded by the
?.and the only surviving evidence is the bare count. There is no capturingILoggeranywhere inDarling.Teststo pass instead.This is the same shape as #2801 and #2816: a failure that reports a plausible-looking value while destroying the one piece of information needed to act on it. A test that can only ever say "expected 17, got 0" cannot be triaged, so it will keep being labelled a flake whether or not it is one.
Suggested fix
Add a minimal capturing
ILoggertoDarling.Testsand pass it at theEnsureRetentionPoliciesAsynccall sites in this test, then include the captured warnings in the assertion message. Then the next occurrence names the actual Postgres error on the first failure instead of costing another round.Worth doing before concluding anything about flakiness, and worth prioritising because this sits on a code path under active change (#2809 retention work, #2811/#2816 fetch-phase work).