Skip to content

Skip locking fresh dimension rows; halve the Query Store liveness-touch rate (#4249, #4250) - #4288

Merged
erikdarlingdata merged 3 commits into
devfrom
fix/4249-dimension-write-volume
Sep 25, 2026
Merged

erikdarlingdata merged 3 commits into
devfrom
fix/4249-dimension-write-volume

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Closes #4249. Part of #4250.

Why

Two sources of needless writes to the Darling store, both measured on production stores:

What changes

#4249: skip fresh digests before any lock

In both branches of UpsertSql (Darling/PerformanceMonitor.Darling.Storage/PayloadDimensions.cs):

  • A pre-filter on the SELECT: WHERE NOT EXISTS (SELECT 1 FROM <dim> d WHERE d.digest = u.digest AND d.last_seen >= $3 - INTERVAL '1 hour'). It is a plain MVCC read, which takes no lock. A digest that is already fresh never reaches INSERT or ON CONFLICT, so it writes nothing.
  • The ON CONFLICT ... WHERE guard stays. It covers the race the pre-filter cannot: two sessions whose pre-filter reads both run before either commits.
  • ORDER BY u.digest on the SELECT. The pre-filter makes the statement a join, and a join can return rows out of input order. A hash anti-join returns them in hash order. The client-side digest sort in PayloadDimensionBatch.ToArrays (Payload-dim batch upserts deadlock under concurrent collection cycles: unnest arrays carry no cross-session lock order #1801) then no longer sets the order in which rows take their locks, and two sessions can deadlock. The ORDER BY sets that order whatever join the planner picks. The client-side sort stays.

Not changed: sending payloads only for new digests saves network traffic, not disk writes, so it is out of scope.

Lite has no dimension tables. It stores query text and plan XML in each row and has no ON CONFLICT upsert. So nothing there locks a row again when it sees the same content.

#4250, items 1 and 2: a 12-hour guard

QueryStoreLivenessTouchGuard.MarginShareDivisor goes from 4 to 2. StampSkewMarginHours stays at 24 (from QueryStorePlanMap.PruneMarginDays = 1), so the guard widens from 6 hours to 12, and the touch rate roughly halves.

  • No prune margin changes. QueryStorePlanMap.MarginOrderingHolds still holds: the map's 1-day margin is still under the dimension margin of 2 days (ChunkIntervalDays + 1).
  • The guard now takes half of the one-day margin, not a quarter. The least room any reachable setting leaves is 2 days (retention floored at 1 day, with the plan-content setting off). 12 hours fits inside that 4 times.
  • A restart does not touch fresh rows again. The guard reads the stored last_seen, not anything the process or the connection holds. A new live test proves it.

Item 3 of #4250, a HOT-only touch, waits for its own measurement.

Measured on a local PostgreSQL 18 rig

This is EXPLAIN (ANALYZE, BUFFERS) of the new statement against a query_plan_dim of 50,000 rows (28 MB). 5,000 rows were marked fresh in the last 5 minutes. The batch had 500 digests: 490 fresh and 10 new.

Insert on query_plan_dim (actual time=1.436..1.437 rows=0.00 loops=1)
  Conflict Resolution: UPDATE
  Conflict Arbiter Indexes: query_plan_dim_pkey
  Tuples Inserted: 10
  Conflicting Tuples: 0
  Buffers: shared hit=2100
  ->  Subquery Scan on "*SELECT*" (actual rows=10.00 loops=1)
        ->  Sort  (Sort Key: u.digest, Sort Method: quicksort  Memory: 28kB) (actual rows=10.00)
              ->  Nested Loop Anti Join (actual rows=10.00 loops=1)
                    ->  Function Scan on u  (unnest, rows=500.00 loops=1)
                    ->  Index Scan using query_plan_dim_pkey on query_plan_dim d
                          Index Cond: (digest = u.digest)
                          Filter: (last_seen >= (now())::timestamp - '01:00:00'::interval)
                          Index Searches: 500
Planning Time: 0.298 ms
Execution Time: 1.524 ms

The planner chose a nested loop anti-join with one index probe per batch row, not a hash anti-join. 10 of the 500 rows passed the pre-filter, 10 were inserted, and none reached ON CONFLICT. It read 2,100 buffers, all from cache, in 1.5 ms. The Sort node under Insert is the ORDER BY. It sets the lock order for either join type.

Tests

  • PayloadDimensionLiveTests.PreFilter_SkipsLockingFreshRows_ButStillRefreshesStaleOnes_AndInsertsNewDigests: a second upsert of the same fresh 50-row batch, 30 minutes later, leaves xmax = 0 on every row, so none was locked. The WAL position moves less than 1 KB. A row stamped 2 hours earlier is refreshed, and a new digest in the same batch is inserted. With the pre-filter removed from one branch, the test failed: all 50 fresh rows came back locked. With the pre-filter restored, it passed.
  • PayloadDimensionTests.UpsertSql_SourceCarriesOrderByOnBothBranches: both branches of UpsertSql carry ORDER BY u.digest.
  • QueryStoreFetchProbeLivePostgresTests.Restart_ANewConnectionRetouchesNothingFresh_TheGuardReadsStoredLastSeen: a plan row and a text row are touched once. Then a new connection probes them again 5 minutes later, inside the guard. No row on the map, plan-dimension or text table is updated.

Test plan

  • dotnet build Darling/Darling.Tests/Darling.Tests.csproj -c Debug: 0 warnings, 0 errors.
  • PayloadDimensionTests, QueryStoreTouchGuardSingleSourceTests, QueryStorePlanFetchTests and DarlingDimensionGcBoundTests: 91 tests, 0 failed. DocCommentHygieneTests: 77 tests, 0 failed.
  • Against the rig: PayloadDimensionLiveTests, 18 tests, 0 failed. QueryStoreFetchProbeLivePostgresTests, 2 tests, 0 failed.
  • Full Darling.Tests once, against the rig on a new darlingtest database: 14,006 tests, 0 failed, 49 skipped, 1 not run. The skips are live classes that need a managed runtime or a log-format cluster. Dev's own CI run shows the same 1 not run.
  • GitHub Actions CI on 53405f7f: build, Darling PostgreSQL tests and Lite tests all pass.

CHANGELOG entry

SECTION: Fixed
ENTRY:

🤖 Generated with Claude Code

https://claude.ai/code/session_01FVjn4PBJN71NQXdFo6ZxNQ

erikdarlingdata and others added 3 commits September 25, 2026 10:12
…touch-guard divisor

PayloadDimensions.UpsertSql: add a WHERE NOT EXISTS pre-filter (plain MVCC read, no
lock) so a digest already fresh within the hour never reaches ON CONFLICT and never
takes a row lock. The ON CONFLICT ... WHERE arm stays for the remaining cross-session
race. Add ORDER BY u.digest to both branches, since the new anti-join lets the planner
return survivors out of unnest's input order; keep the client-side sort too and
rewrite the comment to say why both exist. Update the pinned exact-SQL tests and add a
source pin that both branches carry the ORDER BY.

QueryStoreLivenessTouchGuard.MarginShareDivisor: 4 -> 2 (12-hour guard, half the
touches). No PruneMarginDays changed. Comment rewritten with the measured 6-hour-guard
volume (~163K touches/hour, ~3.9M/day, all non-HOT) and the worst-case headroom (2-day
floor / 12h = 4x).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
PreFilter_SkipsLockingFreshRows_ButStillRefreshesStaleOnes_AndInsertsNewDigests,
against a real Postgres rig:

- A second upsert of the same still-fresh 50-row batch, 30 minutes later, leaves
  xmax = 0 on every row (never locked) and advances pg_current_wal_lsn() by under
  1 KB (measured 0 bytes on a raw-SQL rehearsal of the same shape).
- A row stamped two hours ago is still refreshed.
- A brand-new digest in the same flush is still inserted.

Proved the pin once by reverting the WHERE NOT EXISTS pre-filter: all 50 fresh
rows came back locked (expected 0, actual 50). Restored the fix immediately
after; git diff on PayloadDimensions.cs is empty.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
Restart_ANewConnectionRetouchesNothingFresh_TheGuardReadsStoredLastSeen,
against a real Postgres rig. Stamps a plan and a text row through the touch
guard, then probes the SAME references again from a brand-new connection (the
closest proxy this static, connection-scoped design has to a new writer
instance) five minutes later, still inside the guard window.

Zero rows updated anywhere: the map, plan-dimension, and text relations all
still carry the FIRST touch's last_seen stamp, none carry the second, and the
fresh connection's fetch-probe verdicts still resolve. Confirms the guard's
state lives entirely in the stored last_seen column, not in any writer-side
memory a restart could lose or duplicate.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant