Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion Darling/Darling.Tests/LivePostgresCollectionHygieneTests.cs
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,10 @@ namespace Darling.Tests;
/// <para><b>The exemption is real and must stay available.</b> A class that mints its own database (through
/// <c>ScratchPostgres</c>) or stands up its own cluster does not race the shared store, and serializing it would
/// cost suite time for no safety at all. So this does not demand the attribute — it demands a DECISION, recorded
/// either as the attribute or as an <c>#1776 own-store</c> comment explaining the exemption.</para>
/// either as the attribute or as an <c>#1776 own-store</c> comment explaining the exemption. That isolation is
/// scoped to the rows and relations the own-store class itself created: it does not extend to cluster-wide state
/// such as WAL position, which every session on the same server — own-store scratch databases included — moves
/// (#4354). A test that needs "nothing else touched" must assert against its own rows, not the WAL LSN.</para>
///
/// <para><b>Why the match is the quoted literal and not a substring.</b> <c>DARLING_TEST_PGRUNTIME</c> and
/// <c>DARLING_TEST_PGRUNTIME_OLD</c> are DIFFERENT variables that merely share the prefix, naming an assembled
Expand Down
34 changes: 16 additions & 18 deletions Darling/Darling.Tests/PayloadDimensionLiveTests.cs
Original file line number Diff line number Diff line change
Expand Up @@ -727,12 +727,21 @@ await LiveStoreCleanup.RunAsync(connectionString!, bodySucceeded, async (cleanup
/// against a real server, three ways:
///
/// <para>(1) A second upsert of the SAME still-fresh batch, 30 minutes later, leaves every
/// row's <c>xmax</c> at 0 (never locked) and advances <c>pg_current_wal_lsn()</c> by under 1
/// KB for a 50-row batch — effectively a read-only transaction. This is the pin: reverting the
/// <c>WHERE NOT EXISTS</c> pre-filter (restoring the pre-#4249 shape) turns it red. Confirmed
/// once by hand with a raw-SQL rehearsal of both shapes against this rig: the OLD shape left a
/// real transaction id (781) in <c>xmax</c> on a row whose content and <c>last_seen</c> were
/// both unchanged; the NEW shape left <c>xmax</c> at 0. See PR #4288.</para>
/// row's <c>xmax</c> at 0 (never locked) and its <c>last_seen</c> unchanged — effectively a
/// read-only transaction. This is the pin: reverting the <c>WHERE NOT EXISTS</c> pre-filter
/// (restoring the pre-#4249 shape) turns it red. Confirmed once by hand with a raw-SQL
/// rehearsal of both shapes against this rig: the OLD shape left a real transaction id (781)
/// in <c>xmax</c> on a row whose content and <c>last_seen</c> were both unchanged; the NEW
/// shape left <c>xmax</c> at 0. See PR #4288.
///
/// <para>An earlier revision of this test also asserted a near-zero <c>pg_wal_lsn_diff</c> for
/// the same batch. <c>pg_current_wal_lsn()</c> is cluster-wide, not relation-scoped, so any
/// other session on the same server — the ~39 own-store classes that create scratch databases
/// on this cluster and run in parallel under xunit, plus autovacuum and checkpoints — can push
/// the delta well past the bar even when this test's own transaction touched nothing. The
/// <c>xmax</c> and <c>last_seen</c> checks above already pin "no lock taken, no row touched"
/// directly against the rows this test wrote, with no exposure to unrelated WAL traffic, so
/// the WAL assertion was dropped as redundant and flaky (#4354).</para>
///
/// <para>(2) A row stamped two hours ago is still refreshed — the pre-filter's own staleness
/// read uses the same one-hour boundary the <c>ON CONFLICT ... WHERE</c> guard always used, so
Expand Down Expand Up @@ -807,28 +816,17 @@ async Task<long> ChangedLastSeenCountAsync(byte[][] digests, DateTime expected)
return (long)(await command.ExecuteScalarAsync(ct))!;
}

async Task<string> CurrentWalLsnAsync()
=> (string)(await ScalarAsync(connection, "SELECT pg_current_wal_lsn()::text", ct))!;

/* pg_wal_lsn_diff returns numeric, which Npgsql maps to decimal, not a bigint type. */
async Task<long> WalBytesSinceAsync(string beforeLsn)
=> (long)(decimal)(await ScalarAsync(
connection, "SELECT pg_wal_lsn_diff(pg_current_wal_lsn(), $1::text::pg_lsn)", ct, beforeLsn))!;

// The initial collection cycle: 50 "hot" plans plus one that will go stale.
await FlushAsync(freshPairs.Append((staleDigest, stalePayload)), t0);
Assert.Equal(0, await LockedRowCountAsync(freshDigests));

// (1) Same batch, same digests, 30 minutes later -- still inside the one-hour
// freshness window. The pre-filter excludes every one of them before the statement
// ever reaches INSERT/ON CONFLICT: no lock taken, no last_seen change, near-zero WAL.
var lsnBefore = await CurrentWalLsnAsync();
// ever reaches INSERT/ON CONFLICT: no lock taken, no last_seen change.
await FlushAsync(freshPairs, t0.AddMinutes(30));
var walBytes = await WalBytesSinceAsync(lsnBefore);

Assert.Equal(0, await LockedRowCountAsync(freshDigests));
Assert.Equal(0, await ChangedLastSeenCountAsync(freshDigests, t0));
Assert.True(walBytes < 1024, $"expected a near-zero WAL delta for an all-fresh batch, saw {walBytes} bytes");

// (2) and (3): two hours after t0, the stale row is refreshed and a brand-new digest
// is inserted, in the same flush.
Expand Down
Loading