Repository navigation
Add a v16 managed-conf marker for a 15-minute checkpoint interval (#4246) - #4426
Merged
Merged
Conversation
) Ships checkpoint_timeout = 15min behind a new marker, copying the v15 wal_compression block's shape: builder, marker list entry, append-once heal, and migration classification. The v15 doc comment now says the checkpoint interval ships as v16 rather than being held back. Refs #4246.
erikdarlingdata
marked this pull request as ready for review
September 26, 2026 18:32
This was referenced Sep 26, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Evidence: on the trial store, the 15-minute interval's reading at +3 hours shows WAL volume down from 14.0 to 4.6 GB/h, the full-page-image share down from 8.6% to 4.2%, and the average checkpoint sync phase down from 1.66 s to 0.24 s. That sync phase is far under the self-alert's 10-second bar. The 24-hour reading is logged after merge.
Refs #4246.
Why
The v15 marker shipped
wal_compression = lz4and held back a longer checkpoint interval, because alonger interval risks the checkpointer's own sync-phase self-alert (
DarlingSelfAlertEvaluator.CheckpointSyncBarMs,#4037): more dirty pages per checkpoint makes the fsync phase longer, and #3892 already traced killed
reads on a production store to sync phases of 14.0s and 25.2s. One production store ran a trial of
checkpoint_timeout = 15minthroughALTER SYSTEMand a reload, ahead of this PR, to measure that riskbefore the setting ships to every managed store. Its +3-hour reading is at the top of this description.
What changes
Adds a
v16marker (DarlingManagedPostgres.ConfMarkerV16) that appendscheckpoint_timeout = 15minto
postgresql.conf, in the same shape as the v15 block:BuildCheckpointIntervalConfAppend()builds the marker line plus the setting, with no fingerprint orstamp line (matching v9-v11/v13/v15).
AllManagedConfMarkers, after v15.EnsureConfAppendedappends it after the v15 block, on the same marker-absent guard, with anInformation log line.
ManagedConfMigrationcovers the new marker inCoveredMarkersand classifies its line as a fixedmanaged line (
checkpoint_timeout = 15min).v16.
ManagedConfFile.RenderBodyrenders the v16 block after v15, in the same order as the legacy blocks. A store that has migrated todarling-managed.confrenders only that file and never runs the legacy appenders, so without this line it would never get the setting. A store migrating on first start would carry the value over and then lose it on the next start's render.checkpoint_timeoutissighup-context in PostgreSQL, so a running store could pick it up from areload alone. This service still appends it before
pg_ctl start, the same as every other managedsetting, so a service-owned start applies it on the very start that writes the block.
This block does not touch
max_wal_size:ConfMarkerV12(#3802) already bounds it by free disk, and apin (
WalVolumeConfAppend_PinsV15Marker_AndSetsCompression's v16 counterpart) checks that neithermax_wal_sizenormin_wal_sizenorcheckpoint_completion_targetland in this block by accident.Small stores and size-triggered checkpoints
max_wal_size(ConfMarkerV12, #3802) isclamp(free / 8, 1 GB, 16 GB), floored to a power of two: 1GB, 2 GB, 4 GB, 8 GB, or 16 GB depending on free disk.
PostgreSQL starts a size-triggered checkpoint when WAL written since the last checkpoint approaches
max_wal_size. The PostgreSQL documentation forcheckpoint_completion_targetstates the server aimsto finish each checkpoint's spread-out writes before that much WAL accumulates, which in practice means
a checkpoint fires around
max_wal_size / (1 + checkpoint_completion_target)of WAL, not atmax_wal_sizeitself.ConfMarkerV12pinscheckpoint_completion_target = 0.9on PostgreSQL 14 andlater (the version the managed store ships), so the trigger threshold is
max_wal_size / 1.9.Using one production store's measured post-compression WAL rate of about 4.6 GB/h:
max_wal_sizeWhat a small store gets: at this WAL rate, the 5-minute timer fires before any size trigger on every store, so today every store checkpoints every 5 minutes. With a 15-minute timer, a store whose derived
max_wal_sizeis 1 GB (the smallest, free-disk-limited class) reaches its size threshold first, at about 7 minutes. So it moves from 5 to about 7 minutes, not to 15, and gets a smaller share of the WAL reduction. A 2 GB store moves to about 14 minutes. A store at 4 GB or more gets the full 15 minutes. A store writing less WAL than 4.6 GB/h reaches its size threshold later, so the timer governs more often. Changingmax_wal_sizeis out of scope here: v12 bounds it by free disk on purpose.Compose store
The Linux compose deployment (
Darling/compose/docker-compose.yml) runs the official TimescaleDB image,which runs
timescaledb-tuneon every start. Per #4322, this file already overrides the handful ofsettings tune gets wrong for this product (worker sizing,
work_mem) with explicit flags, and leaveseverything else — including checkpoint settings — to tune's own sizing from the container's cgroup
limit. This PR does not add a checkpoint override to compose: it only ships the v16 marker in the
managed-Windows-store code path, so the compose store's checkpoint interval is unaffected. Whether
timescaledb-tunesetscheckpoint_timeoutitself isn't verified here; if it doesn't, the compose store keeps PostgreSQL's 5-minute default.Test plan
DarlingManagedPostgresTests: addedCheckpointIntervalConfAppend_PinsV16Marker_AndSetsCheckpointTimeout,ConfMarkerV16_IsInAllManagedConfMarkers,EnsureConfAppended_AppendsV16Once_AndNotAgainOnANextStart(copies of the v15 pins), and updated
EveryConfMarker_IsDistinct_AndNoneIsASubstringOfAnother'sexpected count from 15 to 16 markers.
ManagedConfMigrationTests: addedClassifyLines_UntouchedV16Block_IsOurs.ManagedConfFileTests:RenderBody_CarriesEveryManagedConfMarkersOwnedKeys: for every marker inAllManagedConfMarkers, it parses the keys out of that marker's own legacy block and asserts each appears in the rendered managed file, so a future marker with no render line fails it. It fails atc59e1365(before the render line): "marker … (v16 checkpoint interval) owns key 'checkpoint_timeout' … but RenderBody's rendered body does not carry it". It passes with the fix.maintenance_work_memcap, is the one documented skip.RenderBodydecides it from its own inputs, but the memory-sizing cap (2047 MB) sits under the PostgreSQL 17 limit for any RAM size, so no fixture can make the render add it.checkpoint_timeout = '15min'.ManagedConfRehearsalTests: the fixture's managed-key count goes from 24 to 25.ManagedConfFileLiveTests.SecondStart_UnchangedInputs_DoesNotRewriteFile_Gatedcaught the missing render line on CI, and passes with it.DocCommentHygieneTestsrun in-process on macOS: 0 failures introduced (theDarlingManagedPostgresTestsrun shows 5 pre-existing failures, all Windows-path assertions unrelatedto this change, e.g.
ResolveDataDirectory_TrimsATrailingSeparatorexpectingD:\darling\pg).Darling.TestsandLite.Testsboth build 0 errors / 0 warnings-as-errors with-p:EnableWindowsTargeting=true.dev:ConfMarkerV16andBuildCheckpointIntervalConfAppenddon't exist there, so the newtests are a compile failure against the pre-fix code, not a runtime failure.
affect (marker list, migration classifier, doc comments).
CHANGELOG entry
SECTION: Changed
ENTRY:
checkpoint_timeout = 15min, after a trial on one store kept the checkpoint sync phase far under the checkpointer's own sync-phase self-alert.REF:
[Add a v16 managed-conf marker for a 15-minute checkpoint interval (#4246) #4426]: Add a v16 managed-conf marker for a 15-minute checkpoint interval (#4246) #4426