Skip to content

Managed store: cap shared_buffers at 1GB for the co-located reality (v5) + MaxPoolSize=24 — heals the Windows 487 spawn failures - #1559

Merged
erikdarlingdata merged 1 commit into
devfrom
feature/managed-pg-v5-colocated-sizing
Jul 18, 2026
Merged

erikdarlingdata merged 1 commit into
devfrom
feature/managed-pg-v5-colocated-sizing

Conversation

@erikdarlingdata

Copy link
Copy Markdown
Owner

Summary

Round-3 field soak: the .NET-side memory fix holds (service steady ~819MB), but the bundled Postgres kept showing the pre-crash signature — could not reserve shared memory region ... error code 487 on backend spawn, 43 postgres processes, checkpointer working set ~2.5GB. Investigation (sources below): our v3 sizing gave the store a 4GB shared_buffers segment on a 16GB co-located box, and on Windows every backend spawn must re-reserve that segment at the postmaster's base address — with pgsql-bugs history documenting larger shared_buffers exacerbating exactly these 487s.

Changes

  1. DeriveMemorySettings: shared_buffers = min(25% RAM, 1 GB) (was min(25%, 8 GB)). The docs' 25% figure is conditioned on "a dedicated database server" — the managed store is deliberately co-located (service + viewer + often Lite) and double-caches against a shared OS cache. The 1 GB cap shrinks the 487 reservation surface, every backend's reattach, and the checkpointer's sweep.
  2. v5 conf heal block: existing clusters get shared_buffers = <capped> appended under ConfMarkerV5 — postgresql.conf takes the last occurrence, so the override wins without rewriting the v3 block. Fresh installs derive capped in v3 directly; v5 then harmlessly re-states it.
  3. MaxPoolSize = 24 on the service's store connection strings — every pooled Npgsql connection is a live postgres.exe process on Windows, and each spawn crosses the 487 surface; 24 covers the 4-wide sweep + command/beacon/alert/analysis seams.

Field-signal notes (also in the CHANGELOG)

Two signals from the soak were misread in good faith and are documented for the future: Windows charges shared-memory pages to every backend's working set (N processes "totaling" GBs are mostly one segment counted N times), and a ~4.5-minute checkpoint write= phase is deliberate spreading, not distress.

Testing

  • Derivation theory re-pinned across 2/8/16/32/64 GB tiers; v3-block and large-RAM-cap pins updated; new v5 block pins (marker, capped values at two tiers, restates-nothing-else); E2E gains the v5 marker + dedup assertions.
  • Full suites, Release: Darling 2163/0, Lite 1369/0.

Sources: PostgreSQL docs — Resource Consumption (the "dedicated database server" conditioning; no Windows-specific cap exists in current docs — the old one was removed, so the rationale here rests on co-location + the 487 history, not folklore), BUG #14050, BUG #18954.

🤖 Generated with Claude Code

…ol ceiling

The v3 derivation used min(25% RAM, 8GB) - the docs' starting point,
which the docs explicitly condition on a DEDICATED database server.
The managed store is co-located by design (service + viewer + often
Lite), and on Windows the oversize segment is the 487 surface: every
backend spawn re-reserves shared memory at the postmaster's base
address, with pgsql-bugs history (BUG #14050/#18954) documenting
larger shared_buffers exacerbating the failures. Observed live: 16GB
field box, 4GB segment, recurring 487s, checkpointer working set ~=
the buffer pool.

- DeriveMemorySettings: shared_buffers = min(25% RAM, 1GB)
- ConfMarkerV5 heal block: existing clusters get the capped value
  appended (postgresql.conf last-occurrence-wins), no v3 rewrite
- BuildRoleConnectionString: MaxPoolSize=24 bounds the backend count
  (43 processes observed during a 24-server sweep)

Pins updated (theory tiers, v3 8GB example, large-RAM caps) + new v5
block pins + E2E marker assertions. Darling 2163/0, Lite 1369/0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@erikdarlingdata
erikdarlingdata merged commit adc3340 into dev Jul 18, 2026
5 checks passed
@erikdarlingdata
erikdarlingdata deleted the feature/managed-pg-v5-colocated-sizing branch July 18, 2026 00:55
pull Bot pushed a commit to ehtick/PerformanceMonitor that referenced this pull request Jul 29, 2026
…floor (erikdarlingdata#1777)

Field measurement on a production field instance (16 GB RAM class) showed
TimescaleDB compression throughput rising ~70% when maintenance_work_mem went
from the old formula's ~800 MB landing point to 1536 MB, and gaining nothing
measurable at 4096 MB.

The formula becomes min(max(5% RAM, 1536MB), 25% RAM, 2048MB): the 1536 MB
floor is the measured capture point, the 25%-of-RAM term keeps the floor from
overcommitting a small host, and the 2 GB cap bounds the big-RAM case where the
data showed nothing further to gain.

A formula change alone would only ever reach a fresh initdb, and the stores that
need this are already collecting -- so a v7 conf block (ConfMarkerV7) re-states
maintenance_work_mem the same way v5 re-states shared_buffers. postgresql.conf
takes the LAST assignment, so an existing store adopts the raised value on its
next service-owned start without the v3 block ever being rewritten.

Also corrects three stale claims in the Darling README's memory-sizing paragraph
that predate erikdarlingdata#1559 (shared_buffers cap and its 8 GB example, "all three appends").
erikdarlingdata added a commit that referenced this pull request Sep 3, 2026
…2845)

Every conf heal block v1-v7 is keyed on "is this marker absent?", which is
answered once in a cluster's life. Nothing re-derived when the machine was
replaced underneath it, so all three monitoring boxes kept effective_cache_size
at 11.86 GB (75% of 16 GB) after being resized to 31.5 GB.

The v8 block keys on a fingerprint of the derivation inputs instead of a
version, so the two triggers compose: a version marker heals a formula change,
this heals a hardware change. It compares the LAST fingerprint in the conf
rather than any, because postgresql.conf resolves duplicates by
last-occurrence-wins -- a host resized 16 -> 32 -> 16 GB carries both, and a
Contains test would find the stale first one and skip, latching the 32 GB block
in force on a box that no longer has 32 GB.

DeriveWorkerSettings is extracted so v2 and v8 share one worker formula and
cannot drift. Re-stating those counts is also what makes v2's "never goes stale
as collectors are added" claim true; it was true of the formula and false of its
marker-keyed application.

shared_buffers and work_mem are excluded structurally, not by trusting their
caps: shared_buffers because the 1 GB cap is the live Windows error-487
mitigation (#1559) and raising it is a reviewed formula decision rather than
something a resize propagates, work_mem because the formula would double it to
63 MB while every measurement above 31 MB on the store's heaviest read is worse,
and because a per-sort per-connection ceiling follows from the query mix, not
from the machine.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MX6HyjsuDCs15qGB2rh4Gy
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant