Repository navigation
Flaky: DatabaseStateExpectedStoreTests.ResetToCurrent_RebaselinesAndClearsOverride blocked a nightly with no Lite change #2374
Description
Activity
Following the lock lead from the issue body, one concrete thing that is verifiable regardless of what the flake turns out to be.
DuckDbInitializer's lock is process-wide:private static readonly ReaderWriterLockSlim s_dbLock = new(LockRecursionPolicy.NoRecursion);
Every
AcquireReadLock/AcquireWriteLockgoes through that one static, so all 42 test classes serialize against each other on it even though each owns a separate database file. That directly contradicts whatSharedDuckDbFixturebelieves it has bought:Deliberately IClassFixture (one database per class), NOT a collection fixture: a single shared collection would serialize these classes against each other and give back most of the win. Each class owns its own database file, keeping xUnit's cross-class parallelism intact.
The file separation is real, but the parallelism it is protecting is given straight back by the static lock — the classes serialize anyway, just through a lock instead of through a collection, and with no ordering guarantee. That is a plausible contributor to Lite.Tests taking 500 s, and it means cross-class timing is coupled in a way the fixture's reasoning assumes it is not.
Second thing in the same method, which I flag as a hazard rather than a diagnosis:
catch (LockRecursionException) { /* The current thread already owns a read lock — likely leaked by an unhandled exception that prevented Dispose(). ... */ return NoOpDisposable.Instance; }
With
LockRecursionPolicy.NoRecursionthis fires for any nesting on a thread, not only for a leak — a caller that legitimately holds a read lock and calls something that takes another gets an unprotected no-op, and the comment's "we're already protected by a read lock" only holds for the leak case it names. It is a swallow that converts a lock-discipline problem into silent unsynchronized access, which is the shape of failure that shows up as a read returning nothing.I have not tied either of these to the failing assertion and I am not claiming they are the cause. But the static lock in particular is worth fixing on its own terms: the fixture's stated design and the lock's actual scope disagree, and one of the two is wrong.
Second data point, and it confirms the flake: re-dispatched the identical commit as 32285071843 and Lite.Tests passed 2483/2483. Same tree, same test, opposite result — so this is timing, not logic in the change set.
That makes three runs of the same Lite code: pass (16:59), fail (17:47), pass (18:02). The nightly published on the third, so the soak is unblocked, but the failure rate is high enough to bite the release pipeline again.
- added a commit that references this issue
on Aug 19, 2026 Fixed in #2375.
The seed rides the best-effort maintenance block, so a bare sweep can return having written no expectation at all — and the test then asserts on a precondition it never established, failing far from the cause because the deviation read still succeeds with nothing to deviate from.
BaselineAsyncwaits for the baseline to actually land, settling on the recorded state rather than row count (the expectations read LEFT JOINs the snapshot, so an unseeded database is a present row with an empty state).The topology that manufactures this is split out as #2376 — the process-wide store lock hands back the cross-class parallelism
SharedDuckDbFixturebelieves it has. That one is not a release blocker, and the lock must NOT be narrowed: production runs manyDuckDbInitializerinstances against one file and relies on it.No re-cut or re-install: this is test-only, so the soaking artifact (
c78240c5…, built fromc4721626) is still exactly what dev produces.- added a commit that references this issue
on Aug 20, 2026
Lite.Tests.DatabaseStateExpectedStoreTests.ResetToCurrent_RebaselinesAndClearsOverridefailed a nightly build with no Lite code change between a passing and a failing run, so it is intermittent rather than a regression. Filing it rather than re-running quietly, because a flake that blocks the release pipeline gets re-run until it is invisible.Evidence
Between those two runs
devmoved by exactly two commits: #2372 (Darling service + Darling tests) and #2373 (CHANGELOG). Neither touchesLite/orLite.Tests/.Line 223 is the FIRST assertion, before the verb under test ever runs:
So the setup produced no deviation at all: three seeded snapshots (ONLINE at T0, OFFLINE at +1m and +2m) and a baseline read between them, and the deviation query then returned nothing.
Two theories I checked and disproved
Recording these so nobody re-walks them.
Not a wall-clock cliff.
T0is a fixed2026-08-01 09:00:00and today is the 19th, which looks exactly like a fixture aging out of a lookback window. It is not: the onlynow()inLocalDataService.DatabaseStates.csstampsupdated_at, and nothing filtersdatabase_stateson wall clock. A date cliff also cannot produce a pass and a fail on the same day.Not cross-class interference over a shared database. The constructor calls
fixture.ResetData(), and 42 test classes takeSharedDuckDbFixture, which looks like one class wiping another's rows mid-test. It is not: the fixture builds its own database under a GUID temp directory, so each class owns a file, and the type documents that choice deliberately ("Deliberately IClassFixture (one database per class), NOT a collection fixture"). Within a class xUnit is serial.What I have not established
I do not have a root cause and am not going to assert one. What is worth someone's attention:
SeedConnectionAsyncholds one connection per test instance and the reads go throughAcquireReadLock. If a lock acquisition can time out and degrade to an empty result rather than throwing, that would produce this failure signature precisely. I have not read that path closely enough to say.Why it matters beyond one red build
The #2266 scale test is the precedent: it failed on PR after PR whose diffs could not reach it, and the fix was to assert something the product actually owns rather than to keep re-running. Same principle here — an empty collection where a deviation was seeded is either a real ordering bug or a test that does not control its own preconditions, and both are worth knowing before 3.5.0.
Re-dispatched the nightly to unblock the soak artifact. If this reproduces, that is a second data point and it should be fixed rather than retried.