Skip to content

fix(agent/swm): make SWM host-mode store writes crash-safe (atomic temp + fsync + rename) - #2937

Open
branarakic wants to merge 27 commits into
testnet-canaryfrom
fix/swm-host-store-durable-writes
Open

branarakic wants to merge 27 commits into
testnet-canaryfrom
fix/swm-host-store-durable-writes

Conversation

@branarakic

@branarakic branarakic commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

SwmHostModeStore (the per-CG opaque-ciphertext log plus .meta cursor that hosting cores keep under <home>/swm-host/) rewrote its files with a plain fs.writeFile and never fsynced anything. A kill -9 or power loss inside the prune rewrite or a meta write could leave a truncated log or a torn/empty meta, and the "crash between appendFile (durable) and persistMeta (durable)" reasoning in loadMeta rested on writes that were not durable. The prune failure reproduces on the base commit: a child process SIGKILLs itself mid-rewrite and the log is left holding a truncated prefix of the pruned bytes (seqnos 5-6 of 5-9), with the acknowledged entries gone.

User-facing impact. Operators hosting SWM context graphs on cores: a power loss or kill -9 during a prune, a meta write or an append no longer risks a truncated log, a lost cursor, later entries buried behind a torn frame, or a recycled seqno (which breaks host catch-up's strict-greater-than paging), and a log that prune removed because every entry expired stays removed. Cost: each acknowledged append now does three fsyncs (see the decision below), and a prune that removes a whole log does one directory fsync.

Changes

  • Atomic durable rewrites. persistMeta, the loadMeta reconcile write and the prune log rewrite now go through one writeFileDurable: sibling temp file, fsync the temp handle, rename over the target, then fsyncRfc64DirectoryV1 on the directory; the temp file is removed on failure. persistMeta stays authoritative (rejects), the reconcile write stays best-effort, and the all-dropped prune branch removes the log and then directory-syncs the unlink (below).
  • Append fsyncs the frame before the cursor is published (decision below).
  • Leftover temps are inert and swept. A crash leaves <key>.<log|meta>.tmp-<pid>-<uuid>. The existing scans key off the .log / .meta suffix, so they already ignore it, and init() now deletes them: never a temp this same instance still has in flight, never counted as an orphan log, and it does not fire the operator warning on its own (reported as the optional staleTempFilesRemoved).
  • Torn log tail repair (found while writing the crash tests). A partial frame at the tail (crash or ENOSPC mid-append) made every later append invisible to iterate() and let seqnos be reused. Reproduced on the base commit: after a partial frame, two appends were acknowledged as seqno 3 and 4 and iterate() still returned only 1 and 2. The tail is now checked once per CG per process, under the per-CG write lock, and truncated to the last complete frame; it is re-checked after any failed append. Readers already stop at that boundary, so no servable byte is lost. A failed append now burns its seqno instead of risking a duplicate frame.
  • Three hardenings the durable write made necessary.
    1. persistMeta drops the cached meta when its write fails (a retried markRegistered used to no-op against a flag the disk never saw).
    2. loadMeta has one owner per CG: the first caller starts a cold initialization (read .meta, recover the log tail, take the max of the two cursors, best-effort persist the reconciled cursor, install into the cache) and every other caller, unlocked readers and locked mutators alike, awaits that same promise, so no mutation can interleave with it. The earlier cache-install guard only kept a slow cold load from rolling back the in-memory cursor; it did not stop the delayed reconcile rename from overwriting a newer .meta on disk (reproduced: a cleared hostModeSubscribed came back after a restart, and a fresh store re-engaged the graph). The initialization never takes the per-CG write lock (mutators hold it while they await the initialization), a rejected initialization is not cached, persisting the reconciled cursor stays best-effort, and SwmHostModeStore.loadMeta is a non-async delegate returning HostStoreMetaLoader.load's promise (async, body without await), so a non-string id is still a rejection the readers swallow.
    3. A rename or unlink whose directory fsync failed is remembered per change (pendingDirSync, a Map<target, generation>), and an idempotent retry completes that fsync before it acknowledges: markRegistered / markUnregistered / markHostModeSubscribed / markHostModeUnsubscribed whose flag already matches, and a prune that finds the log already pruned or already gone. A directory fsync covers the changes that had returned before it started, so a fsync snapshots the pending (target, generation) pairs first and, once it succeeded, forgets a target only if its entry still has the snapshotted generation: a target that is changed again and fails its own fsync stays pending while an older fsync is in flight. A failing retry keeps the mark, and with nothing pending no extra fsync runs.
  • The all-dropped prune unlink is directory-synced. prune removed a log whose every entry had expired with a bare fs.rm, so a power loss could bring expired ciphertext back. The unlink now goes through the same pending-fsync machinery (one directory fsync per removed log, at prune time only). A cap eviction inside append can take this branch too (a single frame larger than the per-CG byte cap): if the directory fsync after that unlink fails, the append rejects although its frame and cursor were already durable, and the retry stores the next seqno, the same failure mode as a failing rewrite in the eviction path.
  • Module split (file-size ratchet, review round 3). testnet-canary now fails CI when a new source file passes 800 lines or a baselined file grows (node scripts/audit-file-size.mjs). host-mode-store.ts goes from 1,164 to 615 lines; new in packages/agent/src/swm: host-store-durable-fs.ts (242: DurableFiles, atomic replace, durable append/truncate/unlink, pending directory-fsync generations, the live-temp set), host-store-format.ts (158: file naming and frame codec, pure), host-store-meta-loader.ts (152: metadata cache and the one cold-load owner), host-store-startup-reconcile.ts (143: the init sweep), host-store-types.ts (124). The code moved, it was not rewritten: the public types and META_FILE are re-exported from host-mode-store.ts, so no importer changes. cli routes/memory.ts goes from 2,115 to 2,055 lines (base 2,102; its budget in scripts/file-size-baseline.json is lowered to 2,055) by moving the host catch-up route into routes/shared-memory-host-catchup.ts (80). Behavior preservation: a scratch differential harness (not committed) drove the merged head's store and the split one over 42,388 paired runs on a recording in-memory filesystem (647,492 ops, 4,404,439 fs calls per implementation, 214,698 injected failures, a single failure at every call index of 200 sequences), with identical traces, results, errors and final files.
  • New tests, a new devnet suite (devnet/swm-host-store-durability, registered in suites.json, pnpm-workspace.yaml, the root package.json and the lockfile importer), and shared devnet lifecycle helpers. The suite's no-reuse check compares the recovered log with a byte-level snapshot of the frames that were complete at the kill (log-frames.ts, 39 no-devnet tests: no-reuse, cursor coverage with the one-append allowance while the writer runs, served frames, the page-by-page catch-up walk, the quiet-period gate). devnet/_bootstrap/node-lifecycle.ts (63 no-devnet tests; pnpm test:devnet:manifest runs 149) is shared with core-peers-features, which this PR therefore touches: PID-file discovery, liveness, dead-PID cleanup, the kill and restart helpers. Every kill and stop these two suites issue signals only PIDs read from this devnet's own node<N>/{daemon,devnet}.pid files, and only a live PID whose command line is the checkout's CLI entry point (<repo>/packages/cli/dist/cli.js) directly followed by a daemon subcommand as the last argument (a process that merely runs from the checkout, such as the test runner, is refused); a live PID that is not one (a stale file whose number was recycled) makes the helper throw instead of being signalled, and a process that is exiting or is a zombie counts as gone. Restarts go through the same check: restartNodeAndWait stops the node itself (verify every live PID-file entry, SIGTERM, SIGKILL after the grace period, remove dead PID files) before devnet.sh restart-node, whose own stop phase signals PID-file entries unchecked and so finds nothing live. That stop phase also sweeps the process table for processes that mention the node's home directory (its detached children and managed store servers); devnet.sh now limits that sweep to node and oxigraph* processes, so a tail -f <home>/daemon.log (what devnet.sh logs <n> runs), an editor or a grep that only names the home is never signalled. The check reads the checkout path from the command line, not the node's home, because the home (DKG_HOME) is only in the process environment; a recycled PID that became another daemon of the same checkout is not distinguished.

No wire, frame-format, catch-up paging or public-API change. The max(metaSeqno, lastLogSeqno) cold-load reconcile is unchanged. The RFC-64 owner-only file policy is not adopted (atomicity and fsync only); the temp keeps the default mode writeFile used.

Decision: fsync every append (option a)

append fsyncs the frame (appendFileDurable) before persistMeta publishes its seqno, so an acknowledged append is durable in both files.

  • Why not (c), leave the append un-fsynced: because persistMeta is now durable and runs at the end of every append, the seqno cannot be reused after a crash under either choice (a durable cursor ahead of a lost frame just leaves a gap, which strict-greater-than paging tolerates; pinned by a test). What (c) would lose is the frame itself: a hosting core could acknowledge an envelope that a power cut then drops, which is the durability gap this change exists to close, for one more fsync on top of the two that persistMeta already costs.
  • Why not (b), batching: group commit is an API and ordering change in a store whose header says it is deliberately simple. It is the follow-up if throughput ever matters.

Measured, single CG, 500 sequential appends of 4 KiB, on macOS/APFS where Node's fsync is a full flush (an upper bound relative to a Linux SSD):

Variant per append p50 p99
base (no fsync) 0.28 ms 0.16 ms 2.1 ms
(c) durable meta, un-fsynced frame 9.5 ms 9.1 ms 11.3 ms
(a) this PR 16.9 ms 14.9 ms 39.5 ms

The two meta fsyncs the task mandates dominate; (a) adds about 75% on top of (c) on this machine. Appends are already serialized per CG behind the write lock, so one CG's sustained ingest is bounded at roughly 60 per second here (more on a Linux SSD). These figures were measured on the first version of this PR and were not re-measured since: the steady-state append path has not gained an fsync (the pending-fsync bookkeeping is an in-memory lookup, and the retry-after-failure directory fsync only runs after a failure).

Known limits

  • At the byte cap, each append still rewrites the whole surviving log (pre-existing, O(cap) per append), and that rewrite is now fsynced too. At a 64 MiB cap with 256 KiB frames the at-cap p99 was 121 ms vs 154 ms on base (dominated by the page-cache write), so no regression showed, but eviction with a low-water mark would remove the amplification. Not in this PR.
  • The tail repair truncates whatever cannot be parsed as frames at the end of the log. A mid-file corruption that makes a header overrun the file is indistinguishable from a torn tail and is truncated too (those bytes were already unreachable by iterate, stats and prune).
  • fsyncRfc64DirectoryV1 has no soft-fail off Windows, so a filesystem that rejects a directory fsync (EINVAL/ENOTSUP) now makes persistMeta, and therefore append, reject. The RFC-64 durable stores already require the same of the node's data directory.
  • The cold-load and directory-fsync guarantees are per store instance and process (one instance per data dir, as the agent does). A second instance on the same directory has its own lock, initialization and pending marks and is not ordered against the first; opening a fresh instance after the previous one is idle (a restart) is safe.
  • Not tracked: a rename or unlink applied by a process that was killed before its directory fsync (the next process cannot know and takes the visible directory as current; only a power loss inside the kernel's write-back window can still revert it), a rename that rejects yet took effect (taken as atomic), and the unlinks of init()'s sweep of orphan logs, corrupt metas and stale temps (nothing acknowledges them; a resurrected orphan is reaped again by the next init).
  • A load takes only a missing file as absent: any other read error on .meta or the log (EIO, EACCES, EMFILE) rejects the load and the mutation that needed it (nothing cached or persisted), and prune() keeps sweeping past a CG it cannot load and rethrows the first failure.
  • devnet.sh's stop phase still signals a node or oxigraph* process that mentions the node's home directory (for example a DKG_HOME=<home> node cli.js status), as it always has; only other processes that merely name the home are now left alone. The ownership check cannot tell apart a recycled PID that became another daemon of this same checkout, and PIDs are verified once and signalled later (only liveness is re-checked at the signal).
  • Two mutations of the host store survive every suite, as they did before the split: dropping the fsync after the tail-repair truncate, and changing the TTL boundary >= to > in the retention plan.
  • The devnet suite's no-reuse check assumes the retention limits cannot prune during the run (the suite's graph is far below them).
  • No CHANGELOG.md entry: the file has no [Unreleased] section at this base (the top section is the shipped 10.0.21).

Test Plan

Level Tests added or extended Command run Result
Unit packages/agent/test/swm/host-mode-store-{durable-writes,tail-recovery,cold-init,dirsync-retries}.test.ts (new: the former single 64-test durability suite split into four files with a shared test/_helpers/host-mode-store-durability-support.ts, the same 73 test names including the nine unreadable-file tests added since; registered in vitest.unit.config.ts): exact open / fsync / rename / directory-fsync sequence for append, meta and prune; rename, directory-fsync, write, file-fsync and frame-fsync failures; temp cleanup and handle close; cache drop on failure; best-effort reconcile; stale-temp sweep, decoys and the live-temp guard; torn-tail repair (partial payload, partial header, once per process, re-check after a failed append, uninspectable tail); durable cursor ahead of a lost frame; slow cold load vs a concurrent append; readers racing a prune; single-owner cold initialization (a delayed reconcile cannot restore a cleared flag, every mutator and append waits for a paused cold load, concurrent cold loads share one meta read / log scan / reconcile write, two loads in one tick share one initialization, a rejected initialization is retried, a non-string id is a swallowed rejection); retry after a failed directory fsync (table over all four mutators, single and double failure, prune rewrite and prune unlink, reconcile write, covering fsyncs, in-flight and overlapping fsyncs, a target changed again while a covering fsync is in flight, three targets, a cap-eviction unlink). packages/agent/test/swm/host-mode-store-dirsync-model.test.ts (new, 1 test of 16 seeds, about 0.1 to 0.4 s; HOST_STORE_MODEL_SEEDS widens it): a seeded model that holds the mocked directory fsync in flight, settles calls out of order and fails some, and checks every acknowledgement (the four marks, append, prune including unlinks) against a monotonic durable view cd packages/agent && npx vitest run --config vitest.unit.config.ts test/swm/host-mode-store.test.ts test/swm/host-mode-store-durable-writes.test.ts test/swm/host-mode-store-tail-recovery.test.ts test/swm/host-mode-store-cold-init.test.ts test/swm/host-mode-store-dirsync-retries.test.ts test/swm/host-mode-store-dirsync-model.test.ts test/swm/host-mode-store-crash.e2e.test.ts test/swm/host-mode-key-canonicalization.test.ts At this head: 11 files, 168 tests passed (the four split suites, host-mode-store, dirsync-model, key-canonicalization, strip-ciphertext, host-catchup-wire, -requester and -sign). The comparisons with the base below were measured on the earlier single 64-test file and not re-measured against the split suites. Against the pre-PR base host-mode-store.ts (31aff22), 57 of the 64 durability tests (several by timeout, since the base store lacks the features they wait for), the model test and all 9 crash e2e tests fail; the 7 durability tests that pass are deliberate regression pins (clean log untouched, uninspectable tail, lost-frame gap, lagging cursor, a failed initialization not poisoning the queued caller, best-effort reconcile persistence, non-string id). Against the earlier pushed head (6b881be), 34 of 64 and the model test fail. Changed-line coverage at this head: agent 272/276 (98.6%, gate 90%, base a94711b; the uncovered lines are dkg-agent-swm-host.ts:8205, host-mode-store.ts:378, host-store-meta-loader.ts:110, host-store-startup-reconcile.ts:47) on a run of 12 host-store and host catch-up suites, and cli 26/26, computed with the changed-line half of scripts/ci/check-coverage.mjs (inspectCoverage + changedLinesFromDiff + changedCoverage; the script's whole-package floors cannot be met by a partial run, so it was not run end to end). At the round-3 head 28 mutations of the extracted modules (including the four below that the review asked for) each fail tests; one more (the init sweep reaping the log of a .meta it could not read) failed none until host-mode-store.test.ts gained a test. The counts that follow were measured on the earlier single-file version: sharing the cold-load initialization (12 tests fail without it), not caching a rejection (2), installing the cache only after the reconcile write (7), keeping the reconcile write best-effort (3), the log-tail half of the cursor (19), keeping the initialization out of the write lock (behind it, every mutator test deadlocks), the pending-fsync mark and its completion on the early returns (mark*, prune's nothing-to-drop and no-log branches: 23 tests without the mark ever being cleared, 2 without the completion on the no-log return), the per-change generation (the old per-path Set: 2 deterministic tests and the model), and the directory fsync after the prune unlink (6 tests and the model) each fail tests. The earlier-round mutations (removing the tail re-check, the seqno reservation, the live-temp guard, the log fsync, the directory fsync or the temp cleanup) were measured in earlier rounds, against the tests that existed then. Over 2000 seeds the model finds no violation in the current store and hundreds with the per-path Set (372 violations in 259 seeds) or with no fsync after the unlink (997 in 296 seeds); the default 16-seed run caught the per-path Set in 12 of 12 runs the implementer measured and 19 of 20 the reviewer did, so it is probabilistic and the deterministic tests are the gate.
E2E (real files, real process kill) packages/agent/test/swm/host-mode-store-crash.e2e.test.ts (new, 9 tests) and test/_helpers/host-mode-store-crash-child.ts: a tsx child runs a real store and SIGKILLs itself inside prune / persistMeta / append (mid-write, before rename, after rename, torn frame) by wrapping fs.promises; the parent reopens like a restarted daemon and asserts no torn tail, every acknowledged frame served, the cursor never backwards or reused, temps swept, and strict-greater-than paging to the end; every retained envelope is compared byte for byte with the child's fixture after every crash window (store read, paged catch-up read, raw file), and a mid-write / before-rename kill is checked to leave a torn / complete temp same command; and cd packages/agent && npx vitest run --config vitest.config.ts test/swm test/finalized-authority-host-share-paths.test.ts (package full config, every agent SWM suite) 9/9 pass. On the base source 9/9 fail: the prune rewrite leaves a truncated log, an append after a killed mid-frame write leaves a torn tail, and the rename-based kill points are never reached. A prune rewrite that keeps the headers and zero-fills the ciphertext passed the earlier version of this test and now fails the after-rename case. Full-config SWM run at this head: 56 files, 751 tests passed
Devnet devnet/swm-host-store-durability (new suite, registered in suites.json, pnpm-workspace.yaml, the root package.json and the lockfile importer): a 6-node devnet, hosting core node4 and curator edge node5 (config edits backed up and restored), a fresh curated CG; 5 x kill -9 of the core the moment a frame lands, restart, assert on the core's swm-host/ files and on host-mode re-engagement; the curator pages the core with host-catchup to the end from several cursors pnpm test:devnet:swm-host-store-durability and, as regression, pnpm test:devnet:core-peers-features Final live session at 7dadd698e (6 nodes): preflight verified all 12 PIDs of nodes 1-6; swm-host-store-durability 42/42 (3 live + 39 no-devnet, 138 s; 5 kill cycles); core-peers-features 6/6 (63 s, core-fill passing). One change touches the devnet after that session: devnet.sh restart-node's process-table sweep is limited to node and oxigraph* processes (66d411143, and 412f315b7 for a checkout path with a space). Its tests run the real script (a real tail -f <home>/daemon.log survives the stop phase and restartNodeAndWait), but it has not been re-run on a live devnet. The output tail below is from the earlier session, before the module split. Caveat: a process-level kill does not lose page-cache writes, so the suite is a regression and soak check on a real daemon, not a base-versus-fix discriminator; the deterministic base-failing evidence is in the unit and e2e rows. (An earlier version of the suite, with the weaker sorted/unique no-reuse assertion, also passed against the base store build, 3/3 with the base host-mode-store.js swapped into dist; that was not re-run against the current assertions.) The no-reuse check itself was shown to fail on a real daemon: with a dist whose recovery drops the last complete frame on the first append after a restart, the suite fails in cycle 1 ("frame #4 was seqno 4 at the kill but is seqno 5 after recovery"), a log the earlier assertion accepts by construction (run at 7d5e55045; the check has not changed since).
Repo checks pnpm test:devnet:manifest; pnpm run test:scripts; pnpm run lint; pnpm --dir packages/agent exec tsc --noEmit -p tsconfig.json from the worktree manifest 149/149 (5 files; 63 node-lifecycle tests); pnpm test:inventory verified 2265 test files; test:scripts 497/497 pass; node scripts/audit-file-size.mjs passes; lint clean; agent tsc clean; a throwaway tsc over the changed devnet files reports only the pre-existing devnet/_bootstrap/harness.ts(605) HDNodeWallet error, which this PR does not touch

Devnet output

$ pnpm test:devnet:swm-host-store-durability
baseline: host stored 3 frames, last seqno 3
writer shares accepted: 19
cycle 1: meta lags the log (killed between the frame append and the cursor write); at the kill log=4 frames last=4 meta=3; preserved 4 frames byte for byte, 3 new above 4
cycle 2: meta lags the log (killed between the frame append and the cursor write); at the kill log=8 frames last=8 meta=7; preserved 8 frames byte for byte, 2 new above 8
cycle 3: meta lags the log (killed between the frame append and the cursor write); at the kill log=11 frames last=11 meta=10; preserved 11 frames byte for byte, 3 new above 11
cycle 4: temp file left (killed inside a durable meta write); at the kill log=15 frames last=15 meta=14; preserved 15 frames byte for byte, 1 new above 15
cycle 5: meta lags the log (killed between the frame append and the cursor write); at the kill log=17 frames last=17 meta=16; preserved 17 frames byte for byte, 2 new above 17
 ✓ devnet/swm-host-store-durability/log-frames.test.ts (15 tests)
 ✓ the hosting core stores the curator's private SWM shares as opaque frames (non-vacuous gate)
 ✓ 5 x kill -9 during live ingestion: no torn log, no leftover temp, no seqno reuse
 ✓ the curator edge pages the restarted host with strict-greater-than seqnos and reaches the end
 Test Files  2 passed (2)   Tests  18 passed (18)   Duration  229.22s

$ pnpm test:devnet:core-peers-features
 ✓ core-peers-features/automated.test.ts (6 tests) 100223ms
 Test Files  1 passed (1)   Tests  6 passed (6)   Duration  101.58s

Four of the five kills land in the append-to-cursor window (the cursor is one behind the log tail at the kill) and one inside a durable meta write (it left a temp file, which the restart swept). After each restart every frame that was complete at the kill is still in the recovered log, byte for byte, the cursor is recovered from the log tail, and every frame appended after the restart has a seqno above max(cursor, last complete frame) at the kill. Against the base store build the same kills land "between writes" (the base append is near-instant), which is why the suite is a regression check and not a discriminator.

What a devnet core needs before it keeps any private SWM ciphertext in this release (the suite arranges all three and undoes them on exit): the RFC-64 kill switch on the core and the curator (in catalog mode the legacy SWM transport, and host-mode custody with it, is inert); swmHostMode.stripCiphertext=false on the core; and a curated graph that allowlists only the author (a private graph whose roster has a reachable peer besides the author is delivered point-to-point and never gossiped, OT-RFC-49 WS-A). Separately, after the core restarts, the edge-to-core link is relay-only and flaps, so live gossip does not reach the core again until the edge dials it directly; the suite does that after every restart. That last point is an observation about the devnet transport, unrelated to this change.

  • Tests pass locally (targeted lanes above, plus the full-config test/swm run; the full agent lane was not run)
  • Build succeeds (pnpm run build:packages and pnpm --dir packages/cli run build:prepared)
  • Manually tested with dkg start: covered by the devnet suite instead (see the table)

Review follow-up (2026-10-02): a small addition to the catch-up endpoint

Four review findings on the live suite's catch-up scenario needed the endpoint to say more than a count and a cursor. f4d77b1c6 adds two optional body fields to POST /api/shared-memory/host-catchup, the operator's endpoint for confirming what a hosting core holds:

  • maxEntriesPerRound: ask each host for smaller pages. The agent and the wire request already supported it.
  • includeEntries: per peer, the seqno and SHA-256 of every envelope that peer served, in order.

Without them the request and the response are unchanged. 6b3f97b9c then strengthens the scenario: with ingestion stopped the .meta cursor must cover the whole log; paging uses a page a third of the log and must cross page boundaries; and the served envelopes are compared, in order, with the frames on disk (seqno and ciphertext digest). The two new checks are pure functions with no-devnet tests.

At this head: log-frames.test.ts 25/25, the route tests 33/33, agent and CLI type-checks clean, and the live suite 3/3 on a six-node devnet (five kill cycles, each between the frame append and the cursor write; 20 frames served from 0 in four pages).

Review follow-up (2026-10-06): current testnet-canary and the file-size ratchet

The branch merges testnet-canary at a94711bf0. The new file-size audit failed CI on host-mode-store.ts and memory.ts, hence the module split above; the durability suite is split by concern; and the devnet helpers now check ownership on restarts as well as kills (a just-SIGKILLed daemon shows as (node) in state ?E on macOS, which an early version of the verified restart refused; fixed). The last two commits narrow devnet.sh restart-node's own process sweep to node and managed store processes (the executable name decides; a path with a space falls back to what ps -o comm= reports). Not run as one: check-coverage.mjs agent --base end to end (partial runs only, see the unit row).

Related Issues

None filed.

🤖 Generated with Claude Code

Branimir Rakic and others added 5 commits September 30, 2026 05:46
… fsync + rename)

SwmHostModeStore rewrote the per-CG log (prune) and the per-CG meta cursor
with a plain fs.writeFile and never fsynced anything, so a kill -9 or power
loss mid-write could leave a truncated log or a torn/empty meta, and the
"crash between appendFile (durable) and persistMeta (durable)" reasoning in
loadMeta rested on writes that were not durable. A crash inside the prune
rewrite reproduces on the base commit: the log is left holding a truncated
prefix of the pruned bytes and the acknowledged entries are gone.

- Every whole-file rewrite (persistMeta, the loadMeta reconcile write, the
  prune log rewrite) now goes through writeFileDurable: sibling temp file,
  fsync the temp handle, rename over the target, then
  fsyncRfc64DirectoryV1 on the directory. The temp is removed on failure.
  persistMeta stays authoritative (rejects); the reconcile write stays
  best-effort. The all-dropped prune branch still just removes the log.
- Append decision: the frame is now fsynced (appendFileDurable) before the
  cursor is published, so an acknowledged append is durable in both files.
  A failed append burns its seqno instead of risking a duplicate frame.
- Leftover <key>.<log|meta>.tmp-* files from a crash are inert (the scans
  key off the .log/.meta suffix) and are swept by init(); a write in flight
  in the same instance is never reaped.
- Fix found while writing the crash tests: a partial frame at the log tail
  (crash or ENOSPC mid-append) made every later append invisible to
  iterate() and let seqnos be reused. The tail is now checked and truncated
  to the last complete frame once per CG per process, under the write lock,
  and again after a failed append.
- persistMeta drops the cached meta when its write fails, so a retried
  mutation really retries instead of no-op'ing against a flag the disk never
  saw. loadMeta no longer overwrites a cache entry a concurrent locked writer
  installed while its (now durable) reconcile write was in flight.

No wire, frame-format or catch-up paging change.

Tests: unit (host-mode-store-durability), real-file kill -9 e2e with a tsx
child that SIGKILLs itself inside prune / persistMeta / append
(host-mode-store-crash.e2e), both in the agent unit lane.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ng core)

New devnet suite that proves, on a real 6-node devnet, that kill -9 of a core
hosting a curated CG's SWM cannot corrupt SwmHostModeStore's files or recycle
a seqno: the core (node4, stripCiphertext=false) ingests the curator's shares,
is SIGKILLed the moment a new frame lands in its log, restarted, and the
suite asserts on its swm-host/ files (no leftover temp, no torn tail, strictly
increasing seqnos above the pre-kill high-water mark, cursor not below the log
tail, host mode re-engaged) and pages the core from the member with
host-catchup from several sinceSeqno cursors.

Registered in devnet/suites.json, pnpm-workspace.yaml, the root package.json
(test:devnet:swm-host-store-durability) and the lockfile importer.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…up like a member

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…d to end

The hosting core only receives private SWM gossip in this release when the RFC-64
kill switch is on for the core and the curator, swmHostMode.stripCiphertext is
false on the core, and the curated CG allowlists only the author (a roster with
a reachable peer besides the author is delivered point-to-point and never
gossiped). The suite sets those up (backing up and restoring both configs),
dials the curator directly to the core after every restart (the relayed link
flaps), and requests host-catchup from the curator, an allowlisted agent.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@branarakic
branarakic marked this pull request as ready for review September 30, 2026 07:34
@branarakic
branarakic requested a review from Jurij89 as a code owner September 30, 2026 07:34

@otReviewAgent otReviewAgent left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Operational Notice: Review Agent could not complete this review.

Business logic reviewer failed: retry_exhausted

Comment thread packages/agent/src/swm/host-mode-store.ts Outdated
Comment thread packages/agent/src/swm/host-mode-store.ts Outdated
Comment thread devnet/swm-host-store-durability/automated.test.ts Outdated
Comment thread packages/agent/src/swm/host-mode-store.ts Outdated
Comment thread devnet/swm-host-store-durability/automated.test.ts Outdated
Comment thread packages/agent/test/_helpers/host-mode-store-crash-child.ts Outdated
Comment thread packages/agent/test/swm/host-mode-store-crash.e2e.test.ts Outdated
Branimir Rakic and others added 4 commits September 30, 2026 23:00
…d complete a failed directory fsync before acknowledging a retry

Two fixes to packages/agent/src/swm/host-mode-store.ts, found by review of the crash-safe write work, with their tests.

1. Cold metadata initialization

An unlocked cold load (isRegistered / getLastSeqno / stats) prepared a
reconciled snapshot of the .meta and renamed it into place without holding
the per-CG write lock, and a locked mutator that arrived meanwhile ran its own
cold load and its own write. If the reader's rename landed after the
mutation, it restored the older flags on disk. Reproduced deterministically
on the previous head (6b881be): meta seqno 1 + hostModeSubscribed true on
disk, log frames up to seqno 3, an unlocked cold load paused right before its
reconcile rename, markHostModeUnsubscribed completes, the reader is released:
the file says hostModeSubscribed true again and a fresh store re-engages the
subscription. The completion-order "whoever installed the cache first wins"
check only protected the in-memory cursor, never the file.

loadMeta now has exactly one owner per CG. The first caller starts a per-CG
initialization (read .meta, recover the log tail, max of the two cursors,
best-effort persist of the reconciled cursor, install into the cache) and every
other caller awaits the same promise: unlocked readers and locked mutators
alike (append, markRegistered/Unregistered, markHostModeSubscribed/Unsubscribed,
activeLimits from prune and the byte-cap eviction, listHostModeSubscribedCgs,
stats). No mutation can start before the initialization, including its
reconcile write, has finished, so no stale snapshot can be renamed over a newer
meta, and no two callers compute or publish competing snapshots. The
completion-order arbitration in loadMeta is removed as dead code.

Design constraints, checked rather than assumed:
- The initialization does not take the per-CG write lock. Mutators hold that
  lock while they await it, so taking it would deadlock (a scratch mutation that
  wraps the initialization in the lock hangs every mutator test).
- loadMeta stays an async function, with no await before the registration of a
  new initialization: an async body runs synchronously up to its first await, so
  the cache check, the in-flight lookup and the registration are one step that
  nothing can interleave with. Being async also keeps a synchronous throw from
  cgKey() (a non-string id) a rejection that getLastSeqno / isRegistered /
  stats / listHostModeSubscribedCgs swallow through their `.catch(() => null)`,
  as at 6b881be.
- A rejected initialization is not cached: the in-flight entry is dropped when
  it settles and the cache is written only on success, so the next caller
  retries. The cold load still swallows I/O errors exactly as before (an
  unreadable meta or log reads as absent, unchanged); only an unexpected fault
  rejects.
- Persisting the reconciled cursor stays best-effort: a failure there neither
  fails the callers nor poisons the shared initialization.
- Cursor reconciliation is unchanged: max(meta cursor, last complete frame of
  the log), so a seqno is never reused.
- persistMeta's failure path evicts the cache entry only under the per-CG lock,
  and nothing else evicts it, so an eviction cannot land while an
  initialization that was started before it is in flight; the next cold load
  after an eviction starts a fresh initialization from the disk.

Scope of the guarantee: per store instance, per process. The write lock and the
initialization promise are in memory, so two store instances on the same
directory are not ordered against each other. The agent creates one instance
per data dir; opening a fresh instance after the previous one is idle (a
restart, as the tests do) is safe because it re-derives everything from the
files. This is documented in the module header.

Tests (host-mode-store-durability.test.ts, 14 new):
- the interleaving above, asserted on the outcome (disk, then a fresh store)
  and on the mechanism (the mutation waits);
- every mutator plus append arriving during a paused cold load waits for it
  and lands on top of the reconciled cursor (table-driven, 5 rows);
- a mutation that owns the cold load is joined by later readers (one meta
  read, one log scan, two renames);
- concurrent cold loads (4 readers + a mutator) share one initialization (one
  meta read, one log scan, one reconcile write, no re-read once warm);
- two loadMeta calls in the same tick share one initialization (the synchronous
  check-and-register step);
- a rejected initialization is not cached (readers share it, the next caller
  retries) and does not poison the caller queued behind it;
- best-effort reconcile persistence never fails the shared initialization;
- a failed cursor write drops the cache, and the next reader plus a queued
  append share one recovery and stay monotonic;
- a non-string context graph id answers 0 / false from the unlocked readers and
  is skipped by listHostModeSubscribedCgs, as at 6b881be.
The existing "slow unlocked cold load cannot roll the in-memory cursor back"
test pinned the removed arbitration (a concurrent append completing while the
reader is paused); it is adapted to the property that replaces it. It holds the
reader's reconcile rename while the append is in flight and asserts, before
releasing it, that the append has not completed and that only the reader's
(gated) reconcile write has happened, i.e. the append did not start a second
cold load; afterwards both stay monotonic and the reconcile-write count is
exact. (Its first version released the reader before the append could race it
and passed on the previous head; this version is the one that fails there.)

Against 6b881be: 10 of the 12 original cold-initialization tests fail (the
other two are regression pins: a failed initialization does not poison the
queued caller, and best-effort persistence), the adapted slow-cold-load test
fails ("the append must wait for the reader's initialization"), and the
same-tick test fails; the non-string id test passes there, by design.
Scratch mutations of the new code, re-run against the final tests (durability,
store and model files), each fail tests: no sharing (12), rejected initialization
cached (2), cache installed before the reconcile write (7), reconcile
persistence no longer best-effort (3), cursor trusts the meta only (19),
initialization behind the write lock (a deadlock: every mutator test times out,
40 of the 40 selected), an await between the in-flight lookup and its
registration (3: concurrent cold loads, rejected initialization, same tick),
loadMeta not async (1: the non-string id test). An await at the very top of
loadMeta survives, by design: each continuation still checks and registers in
one step.

2. Directory fsync before an idempotent acknowledgement

writeFileDurable renames the new file into place and then fsyncs the
directory. If that fsync throws, the write rejects but the new file is already
visible. persistMeta drops the cache, so the retry of markRegistered /
markUnregistered / markHostModeSubscribed / markHostModeUnsubscribed read the
visible file, found the requested flag already set, and returned success
through the idempotency early return without any directory sync: a power loss
after that acknowledgement could still lose the rename. The prune rewrite has
the same shape: after a rename-then-failed-fsync a retried prune sees the
pruned log as current (nothing left to drop) and returns 0 without syncing. So
does the unlink of a log whose every entry expired, which prune used to do with
a bare fs.rm and never directory-synced at all.

The store now records, per target file, a directory change (a rename over it, or
the unlink of the log) that has not been covered by a successful directory fsync,
and completes that fsync before it acknowledges a no-op on that file:
- pendingDirSync is a Map<target, generation>. A mark is about one change, not
  about a path: syncDirectoryChange(target) takes a generation right after the
  rename or unlink returned and records (target, generation) only when its own
  directory fsync fails. A directory fsync covers exactly the changes that had
  returned before it STARTED, so syncDirectory snapshots the pending
  (target, generation) pairs of the directory before it fsyncs and, once it
  succeeded, deletes a target only if its entry still has the snapshotted
  generation. (The first version of this fix kept a Set of paths. A target that
  was already pending, renamed again while another fsync that had snapshotted it
  was in flight, and failed its own fsync lost its mark when the other fsync
  completed, so the retry was acknowledged with no directory fsync after the
  second rename. The per-path Set is the old semantics; the generation is the
  fix.)
- the four .meta mutators go through one mutateMeta funnel whose early return
  (flag already matches) calls completePendingDirSync first;
- pruneCgUnlocked's "nothing to prune" return and its new "no log" return (the
  retry after a failed unlink fsync) do the same for the log, and the all-dropped
  branch directory-syncs its unlink through syncDirectoryChange;
- a directory fsync that fails keeps the mark and rejects the caller, so a
  second failure is not acknowledged either; any successful directory fsync of
  the same directory (an append's cursor write, another CG's write, the
  cold-load reconcile write) covers older pending changes in it that it had
  seen;
- with nothing pending the early returns do no fsync at all (idempotency after a
  durable success is preserved, pinned by fsync count). The only new fsync on a
  path without a prior failure is one directory fsync per log a prune removes,
  at prune time. The append hot path is unchanged.
The best-effort reconcile write of the cold load participates: if its directory
fsync fails (swallowed), the next idempotent mutation completes it.

Every early return of the store was enumerated: the four mark* mutators and
prune's already-pruned and already-gone branches are the only ones that
acknowledge a state read from a file; append, iterate, stats and the cursor
reads do not acknowledge durability (append always rewrites and syncs).

Behaviour changes to know about:
- A cap eviction inside append can take the all-dropped branch (a single frame
  larger than the per-CG byte cap). If the directory fsync after that unlink
  fails, the append now rejects although its frame and cursor were already
  durable (the frame was evicted at once); the retry stores the next seqno. That
  is the same pre-existing failure mode as a failing rewrite of the log in the
  eviction path.
- The existing test "prune that drops every entry still just removes the log"
  asserted that no durability event happens at all; it now asserts exactly one
  directory fsync (['dirsync:dataDir']) and still no temp file and no rename.
  The other assertions of that test are unchanged.

What is and is not guaranteed, per store instance and process: an acknowledged
.meta write, prune rewrite or prune unlink was covered by a directory fsync
that started after the rename or unlink returned and then succeeded. Not
guaranteed: a rename or unlink applied by a process that was killed before its
directory fsync (the next process cannot know and takes the visible directory
as current; only a power loss inside the kernel's write-back window can still
revert it), a rename that itself rejects (atomic, taken as not applied), the
unacknowledged unlinks of init()'s sweep of orphan logs, corrupt metas and stale
temps (a resurrected orphan is reaped again by the next init), and two store
instances on one directory (they share no lock, no initialization and no pending
marks). The module header says the same.

Tests (host-mode-store-durability.test.ts, 23 new in the directory-fsync groups):
- table over all four mutators: the failed write leaves the requested state
  visible, the retry issues exactly one directory fsync on the data dir, does not
  rewrite the file, and further identical calls issue none;
- table over all four: a second failure keeps the mark, the next retry syncs
  again, then it is free;
- prune retry (single and double failure), and the same for the unlink: the
  retry that finds no log completes the fsync without unlinking again;
- a failed reconcile write's fsync leaves the mark for the next idempotent
  mutation;
- a later successful fsync covers the pending change: an append, another CG's
  meta write, a meta write after a prune rename, an append that recreates the
  unlinked log;
- a rename whose own fsync fails while another fsync is in flight is not cleared
  by it, and the case that broke the per-path Set: a target that is already
  pending, renamed again while the covering fsync is in flight, stays pending
  (with a three-target variant that pins per-target granularity, overlapping
  fsyncs finishing in either order, and one fsync covering every pending target
  while a failing one keeps all of them);
- the cap-eviction variant above.
host-mode-store-dirsync-model.test.ts is a seeded model test of the same
property: the mocked directory fsync is held in flight by a scheduler that
settles calls out of order and fails some (a failing call is quick, a succeeding
one may linger), callers retry or move on, and every acknowledgement (the four
marks, append, prune including the unlinks) is checked against a monotonic
durable view, where a successful fsync makes durable what was visible when it
STARTED and never an older state. It covers five CGs that only flip marks (two
seeds in three) or three CGs that also append past the byte cap and get swept by
concurrent prunes. The default run is 16 seeds (0.1 to 0.4 s on this loaded machine);
HOST_STORE_MODEL_SEEDS widens it. Seeds fix what is drawn, not the interleaving
with real file I/O, so a failing seed points at its printed trace rather than
replaying exactly.

Against the pre-PR base store (31aff22): 57 of the 64 tests of the durability
file, the model test and all 9 crash e2e tests fail (the 7 durability tests that
pass are regression pins: a clean log is untouched, an uninspectable tail is not
appended behind, a failed initialization does not poison the queued caller,
best-effort reconcile persistence, a non-string id, a lost-frame gap, a lagging
cursor). Against the pushed head 6b881be: 34 of 64 and the model test fail.
Scratch mutations each fail tests (the deterministic tests named first; the seeded model test catches several of them only probabilistically, see below): the old per-path Set semantics (2 durability tests), a second failure not refreshing the entry (2), any success clearing every pending target (4), marks cleared before the fsync succeeded (7), a snapshot taken after the fsync (4), every change given the same generation (2), a mark never cleared (23), no fsync after the unlink (6), no completion on the no-log return (2), a failed unlink fsync not remembered (2), the directory fsync before the unlink instead of after (4). The model test is stochastic, so its detection is a rate, not a guarantee: over 2000 seeds it finds no violation in the fixed store and 372 (259 seeds) with the per-path Set, 997 (296 seeds) with no fsync after the unlink, 138 (121 seeds) with it before, and 81 each (47 and 54 seeds) when the retry does not complete it or the failure is not remembered; the default 16-seed run caught the per-path Set in 12 of 12 measured runs by the implementer and in 19 of 20 by the reviewer, and in a full mutation matrix run it sometimes passed for some of these mutants, which is why the deterministic tests are the gate.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… narrow the child's fs interception

The crash e2e (host-mode-store-crash.e2e.test.ts) checked the frames'
seqnos, lengths, cursor and paging after recovery but never their bytes, so a
prune rewrite that kept every header and zero-filled the ciphertext passed all
nine cases (shown: with that scratch mutation of pruneCgUnlocked the previous
test file passes 9/9). The child's fixture (64 bytes, every byte equal to the
seqno) now lives in test/_helpers/host-mode-store-crash-fixture.ts, shared by
the child and the test, and every crash-window recovery compares every retained
envelope byte for byte with it: through the store's whole read, through the
paged catch-up read (pages of 2, strict-greater-than) and on the raw log file
parsed independently of the store (which also covers the frame the parent
appends after recovery). The two prune windows that leave the old log check all
nine frames' bytes right after the kill, and the after-rename prune case checks
seqnos 5..9 on the raw pruned file before restart. With the zero-fill mutation
the after-rename case now fails ("seqno 5 ciphertext differs from what was
written").

The windows are also pinned to where they say they crash, which the old
assertions could not tell apart: a mid-write kill leaves a temp that is a
non-empty strict prefix of the pruned log (prune) or unparseable JSON (meta);
a before-rename kill leaves the complete pruned log (prune) or the complete new
meta (meta) in the temp.

The child (host-mode-store-crash-child.ts) intercepted fs through a generic dispatcher: five write paths
(path-based fs.promises.writeFile and appendFile, handle.writeFile,
handle.appendFile, handle.write) rewritten through `as never`/`Function`
casts, plus a rename wrapper. Running the previous child with each path
labelled showed that across the nine windows only handle.writeFile (prune's
`.log.tmp-*` rewrite and meta's `.meta.tmp-*` write) and handle.appendFile
(append's frame) ever fire; the path-based writes and handle.write never do.
armCrash is now two explicit typed wrappers chosen by the op being crashed,
tearWriteFile (handle.writeFile) and tearAppendFile (handle.appendFile), plus
one rename wrapper; the interception targets are narrowed to the exact file of
each op (the temp sibling for prune and meta), the `armed` flag is gone because
the wrappers are only installed after the setup writes, and the types are the
real fs/promises ones with no casts on the wrapped methods (fs.promises is
patched through a member-typed alias, which type-checks without one). The
kill is still child-local, the write is still torn at half its payload, and
the process still dies by a real SIGKILL.

Same nine crash-window cases before and after (9/9 pass). Scratch mutations
that move a crash point each fail the windows they should: before-rename
crashing after the rename (3 cases), after-rename crashing before it (3),
mid-write tearing the whole payload (3) or zero bytes (3), and an interception
aimed at the wrong file or method (the crash point is never reached: 1, 3, 2,
2 cases).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… fail, share node lifecycle helpers, and refuse to signal a PID that is not a daemon of this checkout

1. The no-reuse check

Devnet suite (swm-host-store-durability): the no-reuse check could not fail on
a recycled seqno. It filtered the recovered log by the pre-kill high-water mark
and then checked that the remaining suffix was sorted. That discards exactly the
frames that would show reuse: with [1,2,3,4] and meta 3 before the kill, a
recovery that dropped complete frame 4 and appended different frames 4 and 5
passed every assertion (sorted, unique, fresh = [5], suffix = [5]) although
seqno 4 was recycled.

The suite now snapshots the complete frames on disk at the kill (seqno plus a
sha256 of each whole frame: timestamp, seqno, length and ciphertext) and the
.meta cursor, and reads them only after the killed processes are really gone
(an unreachable API alone does not say a write already in flight has landed).
After the restart it requires:
  (a) the recovered log starts with exactly that preserved prefix: none of the
      frames complete at the kill dropped, replaced in place or reordered;
  (b) every frame after that prefix has a seqno strictly above the high-water
      mark max(meta cursor, last complete frame) at the kill, strictly
      increasing among themselves, and at least one exists.
The comparison is a pure function in its own module (log-frames.ts:
parseLog, checkNoSeqnoReuse) with no-devnet tests (log-frames.test.ts, 15
cases, added to this suite's vitest config include like the mixed-version and
v10-rs-wallet-rotation suites do for their pure helpers). They include the
assertion-level regression from the review (before [1,2,3,4] with meta 3; a
recovery that drops 4 and appends different 4 and 5 fails, and the legacy
assertion is shown to accept it), the honest recovery (keeps 4, appends 5)
passing, a frame replaced in place, a dropped prefix frame, a dropped tail, a
reordered prefix, a reused seqno above the mark (also when the cursor was
ahead of the log), a duplicate or out-of-order new seqno, and no new frame.
Every existing assertion (no torn tail, no duplicate seqno, sorted, cursor not
below the log tail, temp files swept, host mode re-engaged) is kept. The check
assumes the retention limits cannot prune during the run (the suite's graph is
far below them). Each of seven scratch mutations of the check (digest not
compared, no prefix, cursor ignored, off-by-one at the mark, duplicates
accepted, no-new-frame allowed, dropped tail unreported) fails at least one of
the 15 tests. README updated.

Live evidence at the final head (one devnet session, 6 nodes, 4 cores):
pnpm test:devnet:swm-host-store-durability passed, 3 live tests plus the 15
no-devnet tests (229 s). Every one of the five kills landed in the append ->
cursor write window (four) or inside a durable meta write (one, which left a temp
file that the restart swept), and every cycle reported "preserved N frames byte
for byte, M new above N". Run earlier at 7d5e55045 with the same check: a dist
whose recovery drops the last complete frame fails the live suite in cycle 1
("frame #4 was seqno 4 at the kill but is seqno 5 after recovery"), a log the old
assertion accepts by construction; the check has not changed since.

2. Node lifecycle helpers and the PID guard

The new durability suite carried its own copy of daemon lifecycle knowledge that
core-peers-features already maintained: the PID-file convention, liveness,
dead-PID cleanup, the port environment `scripts/devnet.sh restart-node` needs,
and restart plus readiness. A change to any of these needed two edits. It now
lives in devnet/_bootstrap/node-lifecycle.ts, consumed by both suites; the
crash timing and every SWM / replication assertion stay in the suites.

The module keeps the ownership rules of both suites: a PID is only signalled if
it was read from a PID file of this devnet's own node home, and a PID file is
only removed when its PID is known dead (a live process keeps its file until it
is explicitly stopped). The kill in the durability suite is still an immediate
SIGKILL of every live daemon process the node's PID files list, and
core-peers-features still escalates SIGTERM to SIGKILL in its own
stopNodeProcesses, which stays in the suite and now sits on the shared
primitives.

Where the suites differ, the module takes the choice as an explicit argument
instead of unifying it:
- Hardhat port for restart-node: `rpcUrl` (core-peers-features passes RPC, which
  honours DEVNET_RPC, then node1's chain.rpcUrl, then the default;
  swm-host-store-durability passes node1's chain.rpcUrl only, via
  rpcUrlFromNode1Config). API and libp2p bases still come from node1's config.
- Readiness probe: NodeProbe { authToken, timeoutMs } (core-peers-features sends
  its bearer token with no timeout, polled every 2 s; the durability suite sends
  no token with a 3 s timeout, polled every second).
- An unparseable PID file: core-peers-features leaves it, the durability suite
  removes it with the dead ones; clearDeadNodePidFiles takes removeUnparseable
  (default false) and the durability suite passes true.

Differences that remain and are not byte-for-byte:
- swm-host-store-durability's probe moved from fetch to the plain node:http GET
  core-peers-features has always made (same 200 check, same 3 s abort), so the
  request now carries a JSON Content-Type header on the GET.
- swm-host-store-durability's node1 rpcUrl fallback moved from `??` to a truthy
  check: an empty chain.rpcUrl now falls back to the default instead of throwing
  in `new URL('')`.
- core-peers-features drains the status response instead of accumulating and
  JSON-parsing it (only the status code was ever used), takes waitFor from the
  harness inside the shared wait helpers (same loop and error text as its local
  copy, which it keeps for its own waits), and reads node1's config for the RPC
  through the shared reader with the same precedence (DEVNET_RPC, config,
  default). Its request, headers, intervals and timeouts are unchanged; its kill
  of the victim core now goes through the verified kill below.

A live PID is signalled only if it is a daemon of this checkout. Both previous
copies signalled every live PID read from the PID files, so a stale file whose
number an unrelated process had taken would have had that process SIGKILLed. A
live PID is now signalled only if its command line (`ps -ww -o command=`) is a
DKG daemon started from this checkout: the repo root as a path (as given, or as
its realpath) and a last argument of daemon-supervisor, daemon-worker or
daemon-foreground-worker. A live PID that fails the check, or whose command line
cannot be read, makes verifiedNodePids / sigkillNodeProcesses throw, naming the
file and the PID and never the command line, before anything is signalled: a
silent skip would leave the node running and turn the kill -9 into a graceful
stop. A dead PID is skipped as before.
- Why the checkout and not the node's home: observed with a real
  daemon-supervisor and its daemon-worker, argv is
  `node <repoRoot>/packages/cli/dist/cli.js daemon-supervisor|daemon-worker`;
  the home (DKG_HOME) is only in the process environment. devnet.sh starts nodes
  with DKG_NO_BLUE_GREEN=1, so the entry point is always the checkout's own CLI
  (mixed-version nodes live under <repoRoot>/.devnet-versions). A check on the
  home would need `ps`'s environment output and the exact text devnet.sh
  exported, and a refused live daemon would silently turn a kill -9 into a
  graceful stop.
- What it cannot tell apart: a stale PID recycled by another daemon of this same
  checkout, for example another node of the same devnet. The PID files are then
  the only per-node distinction.
- The durability suite verifies the host's PIDs before its wait for the next
  frame and passes them to sigkillNodeProcesses({ verified }), so the kill at the
  frame stays immediate (the check costs a `ps` per process, tens of
  milliseconds, and the kill window is a few milliseconds wide).
  core-peers-features' stop and victim-kill paths go through the same check.
- A stale claim copied from the base is corrected: devnet.pid is not the
  `cli.js start` launcher that exits after forking. devnet.sh's
  refresh_node_pidfile writes the detached daemon-supervisor (the worker's
  parent) there, or the worker itself when it has been reparented to init.

New no-devnet tests (devnet/_bootstrap/node-lifecycle.test.ts, 40 cases, added
to the include of vitest.manifest.config.ts next to harness-json.test.ts, so
`pnpm test:devnet:manifest` runs them): PID parsing and de-duplication, the
ownership rule with real child processes (a bystander and another node's process
are never signalled, a dead PID in a file is not, a live process keeps its file),
per-node ownership (node homes 1 to 5 exist and node1 holds live decoy PID files:
reading, clearing and killing node2 and node3 see only their own files), the
daemon check (a table of accepted and rejected command lines, the recycled-PID
case with an injected reader and with the real `ps`, both spellings of the root,
the refusal of the verified worker when the other file lists a foreign PID,
no command line in the error), the verified-PID shortcut for time-critical
kills, the two port policies side by side, the readiness probe against a local
server (token sent only when asked, non-200 and refused are false, a hung node
gives up once a timeout is set) and restartNodeAndWait against a stand-in
devnet.sh (args, cwd, port environment, polling, the timeout error, no waiting
when the script fails). Twenty-four scratch mutations of the module, re-run at this head, each fail at
least one of them: among them a nodeHome that resolves every node but 4 to node1
(which survived the first 20 tests), the node number ignored, no command-line
check, any root accepted, the subcommand not required or not last, no path
boundary, an unreadable command line skipped, the verified list ignored, no
realpath spelling, PIDs not de-duplicated, dead PIDs signalled or looked up,
live-PID files deleted, a one-file read, 4xx accepted, the token or timeout
ignored, and the port environment dropped.

Live evidence at the final head (one devnet session): all twelve PID files of the
real 6-node devnet verify (devnet.pid is the daemon-supervisor, reparented to
init; daemon.pid is the daemon-worker whose parent is that supervisor);
core-peers-features passed 6/6 (102 s, core-fill in 42 s) with its stop and
victim-kill paths on the verified kill; the durability suite's five kill cycles
passed with the verified PIDs. Earlier, at 7d5e55045 (the extraction before the
guard), core-peers-features passed 6/6 in both the previous version and the
refactored one, each twice in one session (85 s and 50 s before, 59 s and 86 s
after; no test differed in outcome). The known core-fill flake did not occur in
any run, so nothing is known about its rate here.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…at the review found

The header of devnet/_bootstrap/node-lifecycle.ts now says that devnet.sh honours
a DEVNET_VERSIONS_DIR override (with one outside the checkout a healthy
mixed-version daemon would be refused, loudly; none of the suites here run
mixed-version nodes), and that PIDs are verified once and signalled later with
only a liveness re-check at the signal, so a verified PID that exits and is
recycled inside that window of seconds would still be signalled. Comment text
only.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Comment thread packages/agent/test/swm/host-mode-store-durability.test.ts Outdated
…t-store-durable-writes

# Conflicts:
#	pnpm-workspace.yaml
Comment thread devnet/swm-host-store-durability/automated.test.ts
Comment thread devnet/swm-host-store-durability/automated.test.ts Outdated
Comment thread packages/agent/src/swm/host-mode-store.ts Outdated
branarakic and others added 2 commits October 2, 2026 22:05
POST /api/shared-memory/host-catchup is the operator's way to confirm
what a hosting core holds for a curated graph. It reported counts and a
cursor only, and always asked for the host's default page.

Two optional body fields:
- maxEntriesPerRound asks each host for smaller pages (the agent and
  the wire request already supported it);
- includeEntries adds, per peer, the seqno and SHA-256 of every envelope
  that peer served, in order, whether or not the replay applied it.

Without them the request and the response are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… in host catch-up

Three gaps in the swm-host-store-durability suite's last scenario:

- The kill cycles allow the .meta cursor to trail the log by the frame
  being appended, which also accepts a cursor that is never persisted.
  With ingestion stopped the cursor must now cover the whole log
  (checkCursorCoversLog).
- maxRounds limits requests per call, not entries per page, so paging
  could finish in one data response. The scenario now asks for a page a
  third of the log and requires at least three non-empty pages from 0.
- A matching count and final cursor do not show which frames arrived:
  [1, 1, 3, 4] for a log of [1, 2, 3, 4] has both. The served envelopes
  are now compared, in order, with the frames on disk after the cursor,
  seqno and ciphertext digest (checkServedFrames).

Both checks are pure functions with no-devnet tests for the cases above.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread packages/cli/src/daemon/routes/memory.ts Outdated
Comment thread devnet/swm-host-store-durability/automated.test.ts Outdated
POST /api/shared-memory/host-catchup passes maxEntriesPerRound through to
catchupSwmFromHost. The request digest binds maxEntries, and the requester
signed the value it was asked for while the encoder clamps it to 1024, so a
page size above 1024 produced a signature the host could not verify and the
catch-up fetched nothing. normalizeCatchupMaxEntries applies the encoder's
default and clamp before signing, so the signed and the sent value are the
same. Wire tests: a request signed for 2048 and sent as 1024 fails to verify;
the normalized request verifies at the limit and above it.

The live suite's quiescence gate passed without waiting: its first probe
compared the log with a read taken just before it. quietPeriodGate now
requires the same observation for LOG_QUIET_MS (6 s) and restarts on any
change. It observes the log size together with the .meta cursor, since a
frame is durable before its cursor: a devnet run under load read cursor 2
with three frames on disk in the suite's first test, which now waits on the
same gate before it compares them.

Live run on a six-node devnet: 33/33 (30 unit tests of log-frames, the
baseline, five kill -9 cycles, paging from four cursors).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread packages/agent/test/swm/host-catchup-wire.test.ts
Comment thread devnet/_bootstrap/node-lifecycle.ts Outdated
branarakic and others added 3 commits October 2, 2026 23:19
…sends

The page-size tests in host-catchup-wire.test.ts normalize the page size
themselves, so they stayed green with the normalization removed from
catchupSwmFromHost. The new file runs the real method with a real signer
and a stand-in messenger, decodes what was sent and verifies it as the host
does: for a page above the wire limit, at it, a small one, a fractional one
and the default.

With the previous assignment restored in catchupSwmFromHost, the oversized
case fails with "signer mismatch".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The SWM host store built a context graph's metadata from defaults whenever
the .meta read failed, whatever the error: EIO, EACCES and EMFILE counted as
a missing file. The log-tail recovery did the same for the .log, through an
existence probe and a read that both answered "no log" on any error. The
result was cached, and the next mutation wrote it back over the file.

Since a failed .meta write drops the cache, a warm process can go through
that load again, so one failed write followed by one transient read error
was enough. With the log pruned to empty the cursor restarted at 0: the next
append was acknowledged as seqno 1 and the rewritten .meta no longer carried
hostModeSubscribed, so a requester resuming from the old cursor never saw
the new frames and a restart did not re-engage the subscription. With the
cursor behind the log (the frame of a failed cursor write), a failed log
read let the lagging cursor stand and the next append wrote a seqno that was
already in the log. And a registered context graph loaded as unregistered,
so the next prune applied the unregistered TTL and byte cap to it.

A load now takes only a missing file (ENOENT) as absent: no .meta means
defaults, no .log means no frames. Any other read error on either file
rejects the load and the mutation that needed it; nothing is cached, nothing
is persisted, and the next access reads the files again. The log tail is
recovered by the read alone, without the separate existence probe. The cache
eviction and the failed-mutation rollback are unchanged. So is a .meta that
reads but does not parse: it stays unusable (defaults on load, reaped with
its log by init()). The unlocked readers still swallow the rejection and
answer that one call as if the context graph were unknown. The on-disk
format, the wire format and the method signatures are untouched.

Tests, in host-mode-store-durability.test.ts, eight cases: after a
prune-to-empty and a failed meta write, EIO, EACCES or EMFILE on the re-read
(one case each) rejects the append, writes nothing, and the next append is
seqno 6 with the flag intact across a restart; a .meta that is really
missing on the same route still appends from seqno 1; a transient error is
not cached, the next call on the same store reads again; a reader that hits
the error leaves no defaults behind for the next prune; a failed log read
rejects instead of writing a seqno twice; a failing existence probe cannot
hide the log. With the previous source seven of the eight fail and the
ENOENT control passes: the scenario with 'promise resolved "1" instead of
rejecting', the log-tail case with 'promise resolved "2" instead of
rejecting', the prune case with 63 bytes pruned where 0 were expected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…h it cannot load

With a load that rejects on an unreadable .meta or .log, prune() stopped at
the first context graph whose load failed and left every one after it
unpruned. A context graph that stays unreadable (a damaged .log next to a
readable .meta) would then keep the later ones from ever being pruned.

prune() now skips a context graph whose prune rejects, goes on with the
rest, and rethrows the first failure once the sweep is done, so the prune
timer still logs it.

Test: two context graphs with every frame expired, EIO on the first log
read. prune() rejects with EIO, the other context graph is pruned in that
same sweep, and the next sweep prunes the one that was skipped. Before this
change neither was pruned.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread devnet/_bootstrap/node-lifecycle.ts Outdated
…t-store-durable-writes

# Conflicts:
#	devnet/_bootstrap/vitest.manifest.config.ts
#	devnet/suites.json
#	package.json
#	pnpm-lock.yaml
Comment thread devnet/swm-host-store-durability/automated.test.ts
Branimir Rakic and others added 3 commits October 6, 2026 03:56
…le-size ratchet passes

testnet-canary now fails CI when a baselined file grows past its budget
(node scripts/audit-file-size.mjs). The host catch-up page size and
includeEntries fields (f4d77b1) took packages/cli/src/daemon/routes/memory.ts
from 2,102 to 2,115 lines against a budget of 2,102.

The whole POST /api/shared-memory/host-catchup handler now lives in
packages/cli/src/daemon/routes/shared-memory-host-catchup.ts (80 lines) as
handleSharedMemoryHostCatchupRoute(ctx), which returns false for requests it
does not own; memory.ts dispatches to it where the block stood. The body is
the same: parsing, defaults, the 501 and 500 responses and the response shape
are unchanged (each `return jsonResponse(...)` became `jsonResponse(...);
return true`). memory.ts is 2,055 lines (2,115 before, 2,102 at the base), and
its budget in scripts/file-size-baseline.json is lowered from 2,102 to 2,055
(only that entry: node scripts/audit-file-size.mjs --write would also have
lowered 50 unrelated entries that other merges had already shrunk).

Measured: node scripts/audit-file-size.mjs passes for this file; cli
test/shared-memory-catchup-durable.test.ts (33 tests, which drive the route
through handleMemoryRoutes) and node-ui test/openclaw-bridge.test.ts +
test/ui-compat.test.ts (which read the daemon source) pass; cli tsc --noEmit
is clean.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…mode-store.ts 1,164 -> 615 lines)

testnet-canary's file-size ratchet (node scripts/audit-file-size.mjs) refuses a
new source file over 800 lines, and the host store had reached 1,164 after the
load-error rule and the prune sweep fix. It is split by concern; the code moved,
it was not rewritten:

- host-store-durable-fs.ts (242 lines): DurableFiles, the crash-safe file
  operations (writeFileDurable, unlinkDurable, appendFileDurable,
  truncateFileDurable), the pending directory-fsync bookkeeping
  (pendingDirSync, per-change generations, completePendingDirSync) and the temp
  file lifecycle. Ported from the earlier unpushed round (3facd60fa) onto the
  current store.
- host-store-format.ts (158): file naming (HostStoreLayout) and the frame codec
  (encodeFrame, scanLogFrames, readEntriesSince, countFrames, planRetention,
  concatFrames). Pure functions; every reader keeps the break-on-overrun walk.
- host-store-meta-loader.ts (152): the metadata cache and the one cold-load
  owner per CG (load, initialize, recoverLastSeqnoFromLog), with the
  ENOENT-only rule for a missing file and the no-await registration step kept
  as it was. SwmHostModeStore.loadMeta stays as a non-async delegate returning
  the loader's promise.
- host-store-startup-reconcile.ts (143): reconcileOrphanLogs(dataDir, files),
  the init sweep of orphan logs, corrupt metas and stale temps.
- host-store-types.ts (124): the public types, the default limits and
  CgMetaState. host-mode-store.ts re-exports the public types and META_FILE, so
  no importer changes.
- host-mode-store.ts (615): the public API, the per-CG write lock, sequence
  allocation, the .meta flags, persistMeta, the tail-repair decision and the
  retention policy. Its header lists the modules.

The other agent's changes are intact: a load rejects on any read error but a
missing file (1aec765) and prune() keeps sweeping past a CG it cannot load
(695e68e). Order is untouched: the frame is fsynced before persistMeta, a
failed cursor write drops the cache, an idempotent retry still calls
completePendingDirSync.

Behavior preservation, measured. A scratch differential harness (not committed)
drives the merged head's host-mode-store.ts (1,164 lines, copied) and this tree
over seeded op sequences on an in-memory filesystem that records every fs call
(order and arguments), every file-handle call and every directory fsync, and
completes each call on its own macrotask in a seeded order so concurrent ops
interleave. Sequences mix append, markRegistered/Unregistered,
markHostModeSubscribed/Unsubscribed, prune (with a TTL and a byte cap small
enough to trim and unlink), iterate, stats, the unlocked readers,
init/reconcileOrphanLogsNow, clock jumps and restarts with seeded crash damage
(torn tails, lagging or corrupt metas, orphan logs, stale temps, decoys). 42,388
paired runs (sequential and concurrent; random failures at 2 to 50% of calls;
a single failure injected at every call index of 200 sequences x 2 modes):
647,492 ops, 4,404,439 fs calls per implementation, 214,698 injected
failures. Every trace (calls, completions, results, thrown error codes and
messages, startup reports) and every final file set was identical.

The unchanged host-mode-store, host-mode-store-durability (73 tests),
dirsync-model and key-canonicalization suites pass (112 tests, no test file
edited for this commit); agent tsc --noEmit is clean.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Branimir Rakic and others added 6 commits October 6, 2026 04:29
…wnership before restarts, and unit-test the cursor and paging accounting

Three review findings on the devnet lifecycle helpers, and two gaps found by
comparing the earlier session's unpushed suite changes with what landed.

Ownership check (devnet/_bootstrap/node-lifecycle.ts). isDaemonOfCheckout
accepted any process whose command line held a path under the checkout and
whose last argument was named like a daemon command, so
`node <checkout>/node_modules/vitest/vitest.mjs run daemon-worker` passed and
would have been SIGKILLed from a stale PID file. It now requires the
checkout's CLI entry point (`<root>/packages/cli/dist/cli.js`, as its own
argument, under either spelling of the root) directly followed by
daemon-supervisor, daemon-worker or daemon-foreground-worker as the last
argument, which is what resolveDaemonNodeCommand starts
(`<node> [execArgv] <entry> <subcommand>`). The entry point is searched for in
the `ps` text rather than in whitespace-split tokens, so a checkout path with
spaces still matches. A flag that carries the entry path (`--import=<entry>`)
no longer qualifies; that accept case of the old test is gone.

Restarts (node-lifecycle.ts, swm-host-store-durability, core-peers-features).
`devnet.sh restart-node` stops the node by signalling whatever its PID files
list, with no check, and the suite's setup reached it for node4 and node5
after editing their configs. restartNodeAndWait now runs the new
stopNodeProcesses first: verifiedNodePids (throws before anything is signalled
if a live PID-file entry is not a daemon of this checkout), SIGTERM, SIGKILL
after the grace period, then the dead PID files are removed, so the script's
own stop phase finds no live entry to signal. core-peers-features' private copy
of that stop is replaced by the shared one (same 10 s + 10 s windows). The
swm-host-store-durability setup also verifies both nodes before it edits any
config. Not changed: the script also sweeps the process table for processes
that mention the node's home directory (its managed store servers and detached
children); that sweep is not PID-file based and is documented in the module
header and the suite README.

Node-based API (node-lifecycle.ts). The optional `verified` mode of
sigkillNodeProcesses turned a node-based call into a raw PID-list call; it is
gone. sigkillPids returns the PIDs it signalled, the suite verifies with
verifiedNodePids before its wait and calls sigkillPids at the kill point, and
sigkillNodeProcesses is verify-then-sigkillPids.

Suite assertions (log-frames.ts, automated.test.ts). Ported from the unpushed
round, as the parts the landed versions lack:
- checkCursorCoversLog takes `appendInFlight`: while the writer runs, the
  cursor may trail by the one append in flight, counted in frames. The kill
  cycles used `cursor >= last seqno - 1`, which is measured in seqnos: with a
  seqno burned by a failed append (log 7, 8, 10, cursor 8) it rejects a cursor
  that is exactly one frame behind. The quiescent form is unchanged.
- checkCatchupWalk accounts for a walk page by page (each call resumes at the
  previous cursor, a page holds at most the page size and advances the cursor
  over exactly the frames it served, an empty page ends the walk, exactly
  ceil(frames / page) nonempty pages, and at least 3 from 0), so the claim
  that paging crosses page boundaries is unit-tested instead of living only in
  the live loop; one full response followed by an empty one fails it.
Not ported: the earlier round's per-cycle quiescent re-check (it stops and
restarts the writer in every cycle for a state the final quiescent check
already covers) and its large-share page planner (the landed suite pages with
maxEntriesPerRound).

Tests, measured: node-lifecycle.test.ts 52 tests pass (was 40), including the
vitest example and `ps`-based cases with real child processes (an unrelated
live PID in devnet.pid rejects restartNodeAndWait before devnet.sh runs and
leaves it and the legitimate worker alive; the verified stop signals the
worker and supervisor and leaves bystanders and other nodes alone; SIGKILL
escalation for a process that ignores SIGTERM); log-frames.test.ts 39 tests
pass (was 30). A throwaway tsc over the changed devnet files reports only the
pre-existing devnet/_bootstrap/harness.ts HDNodeWallet error. This commit was
not run against a devnet.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…o four focused files with shared support

Review thread: host-mode-store-durability.test.ts held seven concerns under one
fixture and rebuilt its pause machinery inline. It is now four suites along its
existing describe boundaries, plus a support module (ported from the earlier
unpushed round onto the current file, which the other agent had extended with
the "a file the store could not read is not an absent file" tests):

- host-mode-store-durable-writes.test.ts (324 lines, 14 tests): the atomic
  write sequence and how each step's failure is handled
- host-mode-store-tail-recovery.test.ts (234 lines, 10 tests): leftover temp
  files, torn log tail
- host-mode-store-cold-init.test.ts (547 lines, 26 tests): the single-owner cold
  load, seqno monotonicity across crash windows, and the unreadable-file
  rule (the nine tests the other agent added, moved here)
- host-mode-store-dirsync-retries.test.ts (448 lines, 23 tests): a retry after
  a failed directory fsync, including the unlink cases
- test/_helpers/host-mode-store-durability-support.ts (231 lines):
  useDurabilityFixture (fresh data dir per test, FileHandle prototype, real
  directory fsync, traceDurability, tempFiles, seedLaggingMeta), the frame
  helpers, and the gates createGate, gateRenames ({ every } pauses every
  rename, default only the first) and holdNextDirSync, which replace
  gateFirstRename and the inline copies of it and of holdNextDirSync.

Every suite keeps the outer describe('SwmHostModeStore durable writes') and
declares its own vi.mock of fsyncRfc64DirectoryV1, so each full test name is
unchanged. The four files replace the old one in vitest.unit.config.ts; the
seeded model test's comment points at the retries file.

Preservation, measured: the vitest JSON reporter over the old file (at the
merged head) and over the four new files lists the same 73 full test names
(sorted lists equal), all passing. The 287 lines that carry an assertion are
identical across old and new modulo the `fx.` prefix, except the one line that
races the gate's `reached` promise.

Mutation checks, measured on the extracted modules (src/swm/host-store-*.ts,
host-mode-store.ts), each applied alone to the committed code and run against
the new suites plus the store, model and key-canonicalization suites, against
the old file, and against the differential harness (28 mutations):
- the four mandated ones: no generation check (3 tests: 2 durability, the seeded
  model), no completePendingDirSync (DurableFiles: 19; store call sites: marks 15,
  prune with no log 3, prune with nothing to drop 2), the cursor written before
  the frame (7), the load-error rejection dropped (meta read 5, log read 2);
- the rest of the earlier round's list: no directory fsync after an unlink (7) or a
  rename (29), no temp cleanup (4), no live-temp guard (1), no pending mark (19),
  no temp fsync (6), no frame fsync (2), no tail truncate (4), directory snapshot
  taken after the fsync (5), unshared cold initialization (12), a rejected
  initialization cached (13), the cache installed before the reconcile write (7),
  the reconcile write not best-effort (3), no log tail in the cursor (21), no
  tail re-check after a failed append (1), no seqno reservation (2), the cache
  kept on a failed cursor write (8), no tail repair (4), the prune sweep stopping
  at the first failure (1).
Each fails tests, and the old file and the new files fail the same durability
tests every time. One mutation passed every suite: init()'s sweep reaping a
.log whose .meta it could not read (EMFILE/EACCES), a rule from an earlier
review round that no test pinned; host-mode-store.test.ts now has a test for
it, which fails that mutation (1 test). The differential harness (scratch)
detects 27 of the 28 at its default settings; the one it misses is the missing
generation check, which only the two deterministic overlap tests and the
model catch.

Agent unit lane: the four files (73 tests), host-mode-store (27 tests, one new),
dirsync-model and key-canonicalization pass.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Moving the route into its own module (shared-memory-host-catchup.ts) turned
all of its lines into changed lines for the changed-line coverage gate, and a
partial run (the 33 existing route tests) covered 21 of 26 changed executable
lines, 80.8% against the 80% gate for the cli package. The five uncovered lines
were the error answers, which nothing tested:

- a missing, blank or non-string contextGraphId answers 400 and never reaches
  the agent (4 cases);
- an agent build without catchupSwmFromConnectedHosts answers 501;
- a catch-up that throws answers 500 with the agent's message.

shared-memory-catchup-durable.test.ts has 39 tests (was 33) and passes.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… a restart right after a SIGKILL is not refused

The first live session of the previous commit (6 nodes, real daemons) showed a
regression in it. A preflight over all six nodes passed, and the new suite's
three live tests and the 39 no-devnet tests passed, but its afterAll could not
restore node4 and node5:

  cleanup: could not restore node4: node4: daemon.pid lists pid 64337, which is
  alive but is not a DKG daemon started from <checkout> ...

and core-peers-features, which runs next on the same devnet, then failed at
setup with ECONNREFUSED 127.0.0.1:9204 (node4 was down). The cleanup does
sigkillNode(num), clearDeadNodePidFiles(num), restartNodeAndWait(num). The
SIGKILLed worker is a zombie until its parent (itself killed) or init collects
it: it still answers signal 0, so pidAlive said alive and the PID file stayed,
and `ps` shows it as `(node)`, so restartNodeAndWait's new verified stop
refused it as "not a daemon of this checkout". Before that stop was added, the
shell's restart-node tolerated it. The kill cycles never saw this because they
wait for the killed PIDs to go before restarting.

pidAlive now reads the process state (`ps -o stat=`, new readProcessState) and
counts a zombie as gone; a state that cannot be read counts as alive. That also
makes waitForPidsGone and clearDeadNodePidFiles treat a zombie as gone, and
verifiedNodePids re-checks liveness before it refuses a command line that is not
a daemon's, so a process that exits while it is being looked at is skipped.

Tests (node-lifecycle.test.ts, 57 tests, was 52), each with a real zombie (a
`sleep 0.1 &` whose parent, an `exec sleep`, never waits for it): pidAlive and
readProcessState on a zombie, clearDeadNodePidFiles removing its file,
verifiedNodePids and sigkillNodeProcesses skipping it, stopNodeProcesses next to
a live supervisor, and restartNodeAndWait with a zombie in daemon.pid reaching
devnet.sh. A mutation that makes pidAlive ignore the state fails 5 of them.
pnpm test:devnet:manifest: 143 tests pass (5 files).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…e, not only a zombie

The previous commit counted a zombie (`Z` state) as gone, on the diagnosis that
a SIGKILLed daemon is still a zombie when swm-host-store-durability's afterAll
restarts it. A second live session (6 nodes) failed the same way: node4 and
node5 were not restored, and core-peers-features then failed at setup with
ECONNREFUSED 127.0.0.1:9204. A temporary debug line in verifiedNodePids (not
committed) recorded what it refused: the daemon.pid worker of node4 and of node5
with the command line `(node)` and the `ps` state `?E`, raised from
restartNodeAndWait's stop right after the cleanup's SIGKILL. On macOS the `E`
flag means "the process is trying to exit": it is in the middle of exiting, not
yet a zombie, still answers signal 0 and shows `(node)` instead of its argv.

processStateIsGone(state) is now true for a `Z...` state or one with the `E`
flag, and pidAlive uses it (ordinary states such as `S`, `Ss+`, `R+` and an
unreadable state still count as alive). node-lifecycle.test.ts has a unit test of
the state check with the strings seen on the live devnet, and the real-zombie
tests of the previous commit still apply; 58 tests pass, pnpm test:devnet:manifest
passes. An exiting process cannot be held in that state on demand, so the live
devnet session is what exercises it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…aged store processes

`devnet.sh restart-node` stops a node by scanning `ps eww` for anything that
mentions the node's home directory and signalling every hit (SIGTERM, then
SIGKILL after 15 s) plus its children. A `tail -f <home>/daemon.log` (what
`devnet.sh logs <n>` runs), an editor or a grep that only names the home was
signalled too, so a restart could kill a developer's or another agent's
bystander process on a shared machine, after the TypeScript side had verified
the PID-file entries.

The sweep now keeps a hit only when its executable is `node` or a managed store
binary (`oxigraph*`). PID-file entries and children of kept processes are
unchanged.

Tests run the real script: collect_devnet_node_pids lists a node process and an
`oxigraph-v*` process that mention the home and not a log tail or another node's
process; stop_devnet_node_processes reaps the first two and leaves the tail; and
restartNodeAndWait through the real restart-node (start_node replaced by a
recorder) leaves the tail alive. With the filter reverted three of the four fail.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
* process-table sweep is limited to node and managed store processes.
*/
export async function restartNodeAndWait(paths: DevnetPaths, options: RestartNodeOptions): Promise<void> {
await stopNodeProcesses(paths, options.num, { readCommandLine: options.readCommandLine });

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Issue: Give node stopping one canonical implementation

What's wrong
This adds a second shutdown implementation in front of the existing shell shutdown rather than extending the canonical lifecycle boundary. The two implementations have different process sets, liveness rules, and grace periods, and callers must understand their ordering to understand a restart. The existing restore helper adds another redundant stop.

Example
Restoring an unreachable core now traverses three stop phases: restoreNodeIfDown's stop, restartNodeAndWait's stop, and devnet.sh's stop. The TypeScript and shell implementations separately own discovery, signalling, escalation, waiting, and PID-file cleanup.

Suggested direction
Put ownership verification and shutdown behind one process-control implementation used by both suite helpers and the script. Restart should perform one checked stop, start the node, and probe readiness.

For Agents
Consolidate process control across node-lifecycle.ts and scripts/devnet.sh, then remove redundant stop orchestration from core-peers-features. Preserve ownership verification before signalling, daemon and managed-store cleanup, escalation, and readiness probing. Use the existing lifecycle tests to verify those behaviors.


export class HostStoreMetaLoader {
/** The metadata installed so far, by `cgKey`. The store sets and drops entries as it persists. */
readonly cache = new Map<string, CgMetaState>();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Issue: Keep metadata cache and persistence under one owner

What's wrong
The extraction splits a tightly coupled metadata lifecycle between two classes while exposing its backing Map as their coordination mechanism. The loader owns initialization and reconciliation writes, but the store owns cache replacement, eviction, and authoritative writes. This moves code without establishing a clean ownership boundary, leaving future metadata changes dependent on both classes' internal bookkeeping.

Example
After a metadata write fails, SwmHostModeStore deletes the loader's cache entry. The next load runs reconciliation and a best-effort write inside HostStoreMetaLoader, then a subsequent mutation writes through SwmHostModeStore again. Understanding one metadata retry requires following persistence and cache transitions across both classes.

Suggested direction
Make the cache private and move authoritative metadata persistence and its failure eviction into the metadata component. Expose operations that maintain those invariants instead of exposing the Map.

For Agents
Make the metadata component own loading, persistence, cache installation, and failure invalidation; keep sequence allocation and retention policy in SwmHostModeStore. Preserve burned sequence numbers, shared cold initialization, best-effort reconciliation, and directory-sync retry behavior. Run the existing cold-init and durability retry suites.

…ntains a space

The sweep that limits restart-node's process-table scan to node and managed
store processes judged a process by the first word of its command line. For an
executable under a directory with a space in its name ("My Projects") that word
is a path prefix, so the node's own processes and its oxigraph store would not
have been reaped and the restart would have found the data directory still held.
When the first word does not name a node or oxigraph binary, the executable that
`ps -p <pid> -o comm=` reports decides. A test runs a node whose path has a space
in it; without the fallback it fails.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
expect(log.seqnos[0]).toBe(1);
// The API view agrees with the disk (cleartext or wire-id key).
const stats = await hostStats();
const entries = Object.values(stats?.perCg ?? {}).reduce((sum, row) => sum + row.entries, 0);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Bug: The API consistency assertion can pass using unrelated graphs

What's wrong
The comment claims that the API agrees with the tested graph's disk contents, but the assertion checks only a store-wide minimum. The disk gate remains meaningful; this API check can conceal a missing graph or an incorrect count.

Example
The tested graph has three frames on disk, but the API returns perCg: { unrelated: { entries: 3 } }. This assertion still passes although the tested graph is missing from the API response.

Suggested direction
Select the tested graph's perCg row, require it to exist, and compare its entry count with log.frames.length.

For Agents
Update the baseline stats check in automated.test.ts to resolve the tested graph's cleartext or wire ID and compare its row with the parsed log. Add a negative case where unrelated graphs have enough entries but the tested graph is absent.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants