Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,6 +63,10 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

### Changed

- **Darling store: `maintenance_work_mem` raised to the measured compression floor, and existing stores actually get it** ([#1780], closes [#1777]) - TimescaleDB's compression sort runs on `maintenance_work_mem`, not `work_mem`, so that one setting gates how fast the background compression job moves. Measured on a production field instance (16 GB RAM class), three real points during a one-time catch-up of large backlog chunks: at the old formula's landing point (~800 MB) compression moved **9.1 MB/s** of uncompressed input; at 1536 MB it moved **15.5 MB/s**, with a pure-linear null hypothesis predicting 2833s against an actual 1657s - a real effect, not noise; at 4096 MB it moved 16.1 MB/s, so going 2.7x further past 1536 bought nothing measurable. The formula becomes `min(max(5% RAM, 1536MB), 25% RAM, 2048MB)`. The **1536 MB floor is the fix itself**: the old `min(5% RAM, 1 GB)` landed *under* its own cap on a 16 GB host (819 MB), so raising the cap alone would have changed nothing there. The **25%-of-RAM** term is the small-host guard and the **2 GB cap** bounds the big-RAM case at the point the data stopped improving. Landing values are 1024MB / 1536MB / 1536MB / 1638MB / 2048MB on 4 / 8 / 16 / 32 / 64 GB hosts, each row pinned by a test because each one is a different term winning.

**A formula change alone would have reached only a fresh `initdb`, and every store that needs this is already collecting.** So the change ships with its propagation half: a **v7 marker block** on the same versioned-append pattern v2 through v6 already use, re-stating `maintenance_work_mem` alone the way v5 re-states `shared_buffers`. `postgresql.conf` takes the LAST assignment of a setting, so an existing store adopts the raised value with the v3 block never rewritten - and because the append runs before `pg_ctl start`, it applies on that very start rather than one restart later. Proven end to end against a real PostgreSQL 18.4 + TimescaleDB 2.28.1 store, not asserted: a store was provisioned normally, rewound to its pre-change shape (v7 block removed, v3's value set back to `819MB`), read live at `819MB`, then restarted through the production bootstrap and read live at the new value, with the healed conf showing the legacy line still present and simply outvoted. Every guard was verified RED by mutation: reverting the formula took 10 tests red including every landing row; disabling the heal branch took the propagation E2E red with the existing store keeping 819MB; and dropping the 25% guard took **only** the small-host rows red, leaving 8/16/32/64 GB green, which is the guard proving its own scope. **Honest limits, both carried from the issue:** the three field points are different tables at different sizes rather than a controlled experiment, so the exact threshold between ~800 MB and 1536 MB is unknown and the controlled repro stays open; and the floor raises small hosts proportionally more than the 16 GB class it was measured on (4 GB goes 204 -> 1024 MB). That second one is bounded rather than hand-waved - `maintenance_work_mem` is a per-operation ceiling and not a reservation, PostgreSQL grows the allocation to fit the work, a small host's chunks are small, and PG 17+ (the bundle pins 18.4) made vacuum's dead-TID store grow incrementally instead of allocating the limit up front. One thing operators should not misread: PostgreSQL normalizes units on the way out, so `SHOW maintenance_work_mem` reports `1GB` on a 4 GB host and `2GB` on a 64 GB host. Same setting, different string. That is also why the tests compare bytes via `pg_size_bytes` - the first version compared strings and went red with `Expected: "2048MB", Actual: "2GB"` against a real server

- **CI: a dev/main push now cancels that branch's superseded in-flight build — newest SHA wins** ([#1729]) - [#1715] scoped cancel-in-progress to pull requests and deliberately left every push run uncancellable; 2026-07-26's ~20-merge evening measured what that costs at train speed: of the day's 30 dev-push `build.yml` runs, **13 were superseded mid-flight** (a newer merge landed before the run finished) and **~60 runner-minutes went to SHAs that were already stale** — capacity the shared Windows runner pool bills against every queued PR (#1697's original complaint). Push events in `build.yml` and `sql-validation.yml` now share a per-branch concurrency group with cancel-in-progress, so only the newest head keeps building. "Safe" was verified from the workflow graph, not asserted: a push run produces nothing any other run consumes — every `upload-artifact` in `build.yml` is gated to the release event (the SignPath path) or `failure()` (darling-pg diagnostics), `sql-validation.yml` uploads nothing at all, no `download-artifact` / `gh run download` / `workflow_run` consumer exists anywhere in the repo, nightly builds its own tree from its own checkout, and a release compiles fresh on the `release` event. Release and merge-queue runs keep their unique per-run groups and remain uncancellable: a release waits on SignPath's manual approval gate, and a queue validation is the last check before its result lands on dev. The accepted trade, stated rather than hidden: push builds are diff-scoped, so a cancelled run's areas are not re-verified until the next change touches them — the nightly and the all-areas dev→main release PR are the backstops. The merge queue that would eliminate the train itself stays unavailable on this personal-account repo ([#1716]: organization-owned repositories only, `422 Invalid rule 'merge_queue'`, re-verified against GitHub's GA announcement); this is the slice of that win that IS available here.
- **CI: build.yml handles the merge_group event** ([#1716]) - lands what [#1715] named as the merge-queue prerequisite: without a `merge_group` trigger, the required `build` and `Darling PostgreSQL tests` checks would never report inside a queue and every queued PR would stall. Queue runs take the same always-restore path as dev/main pushes - a queue run is the last validation before its result lands on dev - and the per-run concurrency group, so they are never cancelled or replaced. Path classification works unchanged in a queue: dorny/paths-filter v4.0.1+ resolves merge_group diffs from the payload's base_sha/head_sha whenever the `base` input is empty, which is exactly what the filter steps pass for non-push events. **Post-merge correction to this entry's original "safe one-click" claim: the click does not exist here.** Merge queues are available only on ORGANIZATION-owned repositories (public on any plan, private on Enterprise Cloud - per the GA announcement, and confirmed empirically: creating the ruleset on this personal-account repo returns `422 Invalid rule 'merge_queue'`). The `.gitattributes merge=union` alternative for the CHANGELOG conflict trains is also out: GitHub's server-side PR merging ignores merge attributes (community discussion #9288), so it would only automate local resolution, not the DIRTY state or the re-push. The wiring stays - inert, zero cost, live the day the repo ever moves to an organization - and until then the train-tax relief is [#1715]'s classification fix: single-area re-pushes re-run ~3m40s instead of ~6m30s, docs-only re-pushes 13s.
- **CI: the change classification the v4 pin bump silently broke is restored - and this time it is probe-validated** ([#1715]) - a timing baseline over the 30 most recent build.yml runs plus a throwaway probe PR (#1714) turned the planned build-time audit into a regression find. The 2026-07-26 08:02 pin bump moved dorny/paths-filter v3 -> v4, and v4 evaluates every filter pattern as an INDEPENDENT predicate under its default predicate-quantifier 'some' (a filter is true when any changed file matches at least one rule), so a bare `!**/*.md` exclusion stopped being a subtraction and became its own rule: "any file that is not markdown". Every area filter carrying that line went true for ANY non-markdown change anywhere in the repo - a single root .gitignore edit built and tested all four products (probe run 30219202642), darling-pg ran the full TimescaleDB suite on every PR since the bump including md-only ones (run 30218459544 matched CHANGELOG.md against the darling filter), and the [#1712] docs fast path shipped unable to engage, because every file matches '**' so its code: gate was never false - its measured md-only 1m43s runs were real, but they were the area filters at work (markdown matches no include), not the fast path. The regression was invisible by construction: the bump's own PR touched build.yml, so root=true forced a full build that looks identical to a correct run, and so does every over-built run after it. Fixed keeping v4 (v3 is on the deprecated-runtime track): area filters carry the markdown carve-out INSIDE each include as an extglob (`Darling/**/!(*.md)`) where quantifier semantics cannot detach it; the uninvertible code: filter is replaced by an `all:` counter with docs-only decided by all_count == docs_count; the classify step additionally refuses to engage while any area filter is lit, so the two classifications can never disagree into a `dotnet build --no-restore` with no restore behind it; and check-version-bump.yml, which had the identical '**'-plus-exclusions shape and an equally dead md-only skip, takes the same counter fix. Validated with a per-file truth table on the probe PR (run 30219765613: a Darling .txt probe, a Darling .md probe, .gitignore and the workflow file in one diff - each filter's matched-file list recorded in the PR body). **Also in this pass:** a PR re-push now cancels that PR's superseded in-flight build.yml/sql-validation.yml runs, while push and release runs keep unique per-run concurrency groups - never queued behind or cancelled by anything, every dev/main commit keeps its own check result; the [#1712] allowlist's flagged judgment calls are ratified (CITATION.cff, Screenshots/) but its directory-wide grants become extension-explicit (`docs/**/*.{md,svg,png,jpg,jpeg,gif}`) so a .sql dropped into docs/ tomorrow defaults to code; the scheduled nightly re-dispatches itself onto the dev ref instead of doing real work from main's copy of the workflow file - scheduled workflows execute the DEFAULT branch's copy against dev's checked-out tree, which is exactly how the 2026-07-26 06:00 nightly failed (run 30194606068: main's stale copy read Dashboard/Dashboard.csproj, moved to deprecated/ by #1612 - the #1550 trap again) - so after a ONE-TIME sync of nightly.yml to main (command in the PR body; the schedule stays red each morning until it happens) nightly logic changes take effect the night they merge to dev, with manual dispatches still always building and the artifact job still pinned to dev; and every job that never carries the release/signing path gets a timeout-minutes ceiling at ~3x its worst cold path (darling-pg 30, nightly build 90 / pg 60 / check 10, sql-validation 30 per leg, claude-review 30) so hung-not-slow failures stop holding a shared-pool runner for the 6h default - build.yml's build job stays unbounded on purpose, because the release path waits on SignPath's manual approval gate. **Measured and deliberately not done** (numbers in the PR body): per-area restore splitting (warm restore is 19-33s; four condition-mirrored restore steps buy seconds at the price of the drift risk [#1701] just retired) and cross-job test splitting (Run Lite tests ~2m25s dominates the full build, but a second Windows job costs ~2m45s of checkout/setup/restore/build before its first test - a net loss on a shared serialized pool). Merge queue remains a recommendation with exact settings in the PR body: it is a repo setting, and build.yml needs a merge_group trigger first or queued PRs stall on never-reporting required checks.
Expand Down Expand Up @@ -1805,3 +1809,5 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
[#1773]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/1773
[#1774]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/1774
[#1775]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/1775
[#1777]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/1777
[#1780]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/1780
Loading
Loading