Repository navigation
fix(agent/swm): make SWM host-mode store writes crash-safe (atomic temp + fsync + rename) - #2937
branarakic wants to merge 27 commits into
Conversation
… fsync + rename) SwmHostModeStore rewrote the per-CG log (prune) and the per-CG meta cursor with a plain fs.writeFile and never fsynced anything, so a kill -9 or power loss mid-write could leave a truncated log or a torn/empty meta, and the "crash between appendFile (durable) and persistMeta (durable)" reasoning in loadMeta rested on writes that were not durable. A crash inside the prune rewrite reproduces on the base commit: the log is left holding a truncated prefix of the pruned bytes and the acknowledged entries are gone. - Every whole-file rewrite (persistMeta, the loadMeta reconcile write, the prune log rewrite) now goes through writeFileDurable: sibling temp file, fsync the temp handle, rename over the target, then fsyncRfc64DirectoryV1 on the directory. The temp is removed on failure. persistMeta stays authoritative (rejects); the reconcile write stays best-effort. The all-dropped prune branch still just removes the log. - Append decision: the frame is now fsynced (appendFileDurable) before the cursor is published, so an acknowledged append is durable in both files. A failed append burns its seqno instead of risking a duplicate frame. - Leftover <key>.<log|meta>.tmp-* files from a crash are inert (the scans key off the .log/.meta suffix) and are swept by init(); a write in flight in the same instance is never reaped. - Fix found while writing the crash tests: a partial frame at the log tail (crash or ENOSPC mid-append) made every later append invisible to iterate() and let seqnos be reused. The tail is now checked and truncated to the last complete frame once per CG per process, under the write lock, and again after a failed append. - persistMeta drops the cached meta when its write fails, so a retried mutation really retries instead of no-op'ing against a flag the disk never saw. loadMeta no longer overwrites a cache entry a concurrent locked writer installed while its (now durable) reconcile write was in flight. No wire, frame-format or catch-up paging change. Tests: unit (host-mode-store-durability), real-file kill -9 e2e with a tsx child that SIGKILLs itself inside prune / persistMeta / append (host-mode-store-crash.e2e), both in the agent unit lane. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ng core) New devnet suite that proves, on a real 6-node devnet, that kill -9 of a core hosting a curated CG's SWM cannot corrupt SwmHostModeStore's files or recycle a seqno: the core (node4, stripCiphertext=false) ingests the curator's shares, is SIGKILLed the moment a new frame lands in its log, restarted, and the suite asserts on its swm-host/ files (no leftover temp, no torn tail, strictly increasing seqnos above the pre-kill high-water mark, cursor not below the log tail, host mode re-engaged) and pages the core from the member with host-catchup from several sinceSeqno cursors. Registered in devnet/suites.json, pnpm-workspace.yaml, the root package.json (test:devnet:swm-host-store-durability) and the lockfile importer. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…up like a member Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…d to end The hosting core only receives private SWM gossip in this release when the RFC-64 kill switch is on for the core and the curator, swmHostMode.stripCiphertext is false on the core, and the curated CG allowlists only the author (a roster with a reachable peer besides the author is delivered point-to-point and never gossiped). The suite sets those up (backing up and restoring both configs), dials the curator directly to the core after every restart (the relayed link flaps), and requests host-catchup from the curator, an allowlisted agent. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
otReviewAgent
left a comment
There was a problem hiding this comment.
Operational Notice: Review Agent could not complete this review.
Business logic reviewer failed: retry_exhausted
…d complete a failed directory fsync before acknowledging a retry Two fixes to packages/agent/src/swm/host-mode-store.ts, found by review of the crash-safe write work, with their tests. 1. Cold metadata initialization An unlocked cold load (isRegistered / getLastSeqno / stats) prepared a reconciled snapshot of the .meta and renamed it into place without holding the per-CG write lock, and a locked mutator that arrived meanwhile ran its own cold load and its own write. If the reader's rename landed after the mutation, it restored the older flags on disk. Reproduced deterministically on the previous head (6b881be): meta seqno 1 + hostModeSubscribed true on disk, log frames up to seqno 3, an unlocked cold load paused right before its reconcile rename, markHostModeUnsubscribed completes, the reader is released: the file says hostModeSubscribed true again and a fresh store re-engages the subscription. The completion-order "whoever installed the cache first wins" check only protected the in-memory cursor, never the file. loadMeta now has exactly one owner per CG. The first caller starts a per-CG initialization (read .meta, recover the log tail, max of the two cursors, best-effort persist of the reconciled cursor, install into the cache) and every other caller awaits the same promise: unlocked readers and locked mutators alike (append, markRegistered/Unregistered, markHostModeSubscribed/Unsubscribed, activeLimits from prune and the byte-cap eviction, listHostModeSubscribedCgs, stats). No mutation can start before the initialization, including its reconcile write, has finished, so no stale snapshot can be renamed over a newer meta, and no two callers compute or publish competing snapshots. The completion-order arbitration in loadMeta is removed as dead code. Design constraints, checked rather than assumed: - The initialization does not take the per-CG write lock. Mutators hold that lock while they await it, so taking it would deadlock (a scratch mutation that wraps the initialization in the lock hangs every mutator test). - loadMeta stays an async function, with no await before the registration of a new initialization: an async body runs synchronously up to its first await, so the cache check, the in-flight lookup and the registration are one step that nothing can interleave with. Being async also keeps a synchronous throw from cgKey() (a non-string id) a rejection that getLastSeqno / isRegistered / stats / listHostModeSubscribedCgs swallow through their `.catch(() => null)`, as at 6b881be. - A rejected initialization is not cached: the in-flight entry is dropped when it settles and the cache is written only on success, so the next caller retries. The cold load still swallows I/O errors exactly as before (an unreadable meta or log reads as absent, unchanged); only an unexpected fault rejects. - Persisting the reconciled cursor stays best-effort: a failure there neither fails the callers nor poisons the shared initialization. - Cursor reconciliation is unchanged: max(meta cursor, last complete frame of the log), so a seqno is never reused. - persistMeta's failure path evicts the cache entry only under the per-CG lock, and nothing else evicts it, so an eviction cannot land while an initialization that was started before it is in flight; the next cold load after an eviction starts a fresh initialization from the disk. Scope of the guarantee: per store instance, per process. The write lock and the initialization promise are in memory, so two store instances on the same directory are not ordered against each other. The agent creates one instance per data dir; opening a fresh instance after the previous one is idle (a restart, as the tests do) is safe because it re-derives everything from the files. This is documented in the module header. Tests (host-mode-store-durability.test.ts, 14 new): - the interleaving above, asserted on the outcome (disk, then a fresh store) and on the mechanism (the mutation waits); - every mutator plus append arriving during a paused cold load waits for it and lands on top of the reconciled cursor (table-driven, 5 rows); - a mutation that owns the cold load is joined by later readers (one meta read, one log scan, two renames); - concurrent cold loads (4 readers + a mutator) share one initialization (one meta read, one log scan, one reconcile write, no re-read once warm); - two loadMeta calls in the same tick share one initialization (the synchronous check-and-register step); - a rejected initialization is not cached (readers share it, the next caller retries) and does not poison the caller queued behind it; - best-effort reconcile persistence never fails the shared initialization; - a failed cursor write drops the cache, and the next reader plus a queued append share one recovery and stay monotonic; - a non-string context graph id answers 0 / false from the unlocked readers and is skipped by listHostModeSubscribedCgs, as at 6b881be. The existing "slow unlocked cold load cannot roll the in-memory cursor back" test pinned the removed arbitration (a concurrent append completing while the reader is paused); it is adapted to the property that replaces it. It holds the reader's reconcile rename while the append is in flight and asserts, before releasing it, that the append has not completed and that only the reader's (gated) reconcile write has happened, i.e. the append did not start a second cold load; afterwards both stay monotonic and the reconcile-write count is exact. (Its first version released the reader before the append could race it and passed on the previous head; this version is the one that fails there.) Against 6b881be: 10 of the 12 original cold-initialization tests fail (the other two are regression pins: a failed initialization does not poison the queued caller, and best-effort persistence), the adapted slow-cold-load test fails ("the append must wait for the reader's initialization"), and the same-tick test fails; the non-string id test passes there, by design. Scratch mutations of the new code, re-run against the final tests (durability, store and model files), each fail tests: no sharing (12), rejected initialization cached (2), cache installed before the reconcile write (7), reconcile persistence no longer best-effort (3), cursor trusts the meta only (19), initialization behind the write lock (a deadlock: every mutator test times out, 40 of the 40 selected), an await between the in-flight lookup and its registration (3: concurrent cold loads, rejected initialization, same tick), loadMeta not async (1: the non-string id test). An await at the very top of loadMeta survives, by design: each continuation still checks and registers in one step. 2. Directory fsync before an idempotent acknowledgement writeFileDurable renames the new file into place and then fsyncs the directory. If that fsync throws, the write rejects but the new file is already visible. persistMeta drops the cache, so the retry of markRegistered / markUnregistered / markHostModeSubscribed / markHostModeUnsubscribed read the visible file, found the requested flag already set, and returned success through the idempotency early return without any directory sync: a power loss after that acknowledgement could still lose the rename. The prune rewrite has the same shape: after a rename-then-failed-fsync a retried prune sees the pruned log as current (nothing left to drop) and returns 0 without syncing. So does the unlink of a log whose every entry expired, which prune used to do with a bare fs.rm and never directory-synced at all. The store now records, per target file, a directory change (a rename over it, or the unlink of the log) that has not been covered by a successful directory fsync, and completes that fsync before it acknowledges a no-op on that file: - pendingDirSync is a Map<target, generation>. A mark is about one change, not about a path: syncDirectoryChange(target) takes a generation right after the rename or unlink returned and records (target, generation) only when its own directory fsync fails. A directory fsync covers exactly the changes that had returned before it STARTED, so syncDirectory snapshots the pending (target, generation) pairs of the directory before it fsyncs and, once it succeeded, deletes a target only if its entry still has the snapshotted generation. (The first version of this fix kept a Set of paths. A target that was already pending, renamed again while another fsync that had snapshotted it was in flight, and failed its own fsync lost its mark when the other fsync completed, so the retry was acknowledged with no directory fsync after the second rename. The per-path Set is the old semantics; the generation is the fix.) - the four .meta mutators go through one mutateMeta funnel whose early return (flag already matches) calls completePendingDirSync first; - pruneCgUnlocked's "nothing to prune" return and its new "no log" return (the retry after a failed unlink fsync) do the same for the log, and the all-dropped branch directory-syncs its unlink through syncDirectoryChange; - a directory fsync that fails keeps the mark and rejects the caller, so a second failure is not acknowledged either; any successful directory fsync of the same directory (an append's cursor write, another CG's write, the cold-load reconcile write) covers older pending changes in it that it had seen; - with nothing pending the early returns do no fsync at all (idempotency after a durable success is preserved, pinned by fsync count). The only new fsync on a path without a prior failure is one directory fsync per log a prune removes, at prune time. The append hot path is unchanged. The best-effort reconcile write of the cold load participates: if its directory fsync fails (swallowed), the next idempotent mutation completes it. Every early return of the store was enumerated: the four mark* mutators and prune's already-pruned and already-gone branches are the only ones that acknowledge a state read from a file; append, iterate, stats and the cursor reads do not acknowledge durability (append always rewrites and syncs). Behaviour changes to know about: - A cap eviction inside append can take the all-dropped branch (a single frame larger than the per-CG byte cap). If the directory fsync after that unlink fails, the append now rejects although its frame and cursor were already durable (the frame was evicted at once); the retry stores the next seqno. That is the same pre-existing failure mode as a failing rewrite of the log in the eviction path. - The existing test "prune that drops every entry still just removes the log" asserted that no durability event happens at all; it now asserts exactly one directory fsync (['dirsync:dataDir']) and still no temp file and no rename. The other assertions of that test are unchanged. What is and is not guaranteed, per store instance and process: an acknowledged .meta write, prune rewrite or prune unlink was covered by a directory fsync that started after the rename or unlink returned and then succeeded. Not guaranteed: a rename or unlink applied by a process that was killed before its directory fsync (the next process cannot know and takes the visible directory as current; only a power loss inside the kernel's write-back window can still revert it), a rename that itself rejects (atomic, taken as not applied), the unacknowledged unlinks of init()'s sweep of orphan logs, corrupt metas and stale temps (a resurrected orphan is reaped again by the next init), and two store instances on one directory (they share no lock, no initialization and no pending marks). The module header says the same. Tests (host-mode-store-durability.test.ts, 23 new in the directory-fsync groups): - table over all four mutators: the failed write leaves the requested state visible, the retry issues exactly one directory fsync on the data dir, does not rewrite the file, and further identical calls issue none; - table over all four: a second failure keeps the mark, the next retry syncs again, then it is free; - prune retry (single and double failure), and the same for the unlink: the retry that finds no log completes the fsync without unlinking again; - a failed reconcile write's fsync leaves the mark for the next idempotent mutation; - a later successful fsync covers the pending change: an append, another CG's meta write, a meta write after a prune rename, an append that recreates the unlinked log; - a rename whose own fsync fails while another fsync is in flight is not cleared by it, and the case that broke the per-path Set: a target that is already pending, renamed again while the covering fsync is in flight, stays pending (with a three-target variant that pins per-target granularity, overlapping fsyncs finishing in either order, and one fsync covering every pending target while a failing one keeps all of them); - the cap-eviction variant above. host-mode-store-dirsync-model.test.ts is a seeded model test of the same property: the mocked directory fsync is held in flight by a scheduler that settles calls out of order and fails some (a failing call is quick, a succeeding one may linger), callers retry or move on, and every acknowledgement (the four marks, append, prune including the unlinks) is checked against a monotonic durable view, where a successful fsync makes durable what was visible when it STARTED and never an older state. It covers five CGs that only flip marks (two seeds in three) or three CGs that also append past the byte cap and get swept by concurrent prunes. The default run is 16 seeds (0.1 to 0.4 s on this loaded machine); HOST_STORE_MODEL_SEEDS widens it. Seeds fix what is drawn, not the interleaving with real file I/O, so a failing seed points at its printed trace rather than replaying exactly. Against the pre-PR base store (31aff22): 57 of the 64 tests of the durability file, the model test and all 9 crash e2e tests fail (the 7 durability tests that pass are regression pins: a clean log is untouched, an uninspectable tail is not appended behind, a failed initialization does not poison the queued caller, best-effort reconcile persistence, a non-string id, a lost-frame gap, a lagging cursor). Against the pushed head 6b881be: 34 of 64 and the model test fail. Scratch mutations each fail tests (the deterministic tests named first; the seeded model test catches several of them only probabilistically, see below): the old per-path Set semantics (2 durability tests), a second failure not refreshing the entry (2), any success clearing every pending target (4), marks cleared before the fsync succeeded (7), a snapshot taken after the fsync (4), every change given the same generation (2), a mark never cleared (23), no fsync after the unlink (6), no completion on the no-log return (2), a failed unlink fsync not remembered (2), the directory fsync before the unlink instead of after (4). The model test is stochastic, so its detection is a rate, not a guarantee: over 2000 seeds it finds no violation in the fixed store and 372 (259 seeds) with the per-path Set, 997 (296 seeds) with no fsync after the unlink, 138 (121 seeds) with it before, and 81 each (47 and 54 seeds) when the retry does not complete it or the failure is not remembered; the default 16-seed run caught the per-path Set in 12 of 12 measured runs by the implementer and in 19 of 20 by the reviewer, and in a full mutation matrix run it sometimes passed for some of these mutants, which is why the deterministic tests are the gate. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… narrow the child's fs interception
The crash e2e (host-mode-store-crash.e2e.test.ts) checked the frames'
seqnos, lengths, cursor and paging after recovery but never their bytes, so a
prune rewrite that kept every header and zero-filled the ciphertext passed all
nine cases (shown: with that scratch mutation of pruneCgUnlocked the previous
test file passes 9/9). The child's fixture (64 bytes, every byte equal to the
seqno) now lives in test/_helpers/host-mode-store-crash-fixture.ts, shared by
the child and the test, and every crash-window recovery compares every retained
envelope byte for byte with it: through the store's whole read, through the
paged catch-up read (pages of 2, strict-greater-than) and on the raw log file
parsed independently of the store (which also covers the frame the parent
appends after recovery). The two prune windows that leave the old log check all
nine frames' bytes right after the kill, and the after-rename prune case checks
seqnos 5..9 on the raw pruned file before restart. With the zero-fill mutation
the after-rename case now fails ("seqno 5 ciphertext differs from what was
written").
The windows are also pinned to where they say they crash, which the old
assertions could not tell apart: a mid-write kill leaves a temp that is a
non-empty strict prefix of the pruned log (prune) or unparseable JSON (meta);
a before-rename kill leaves the complete pruned log (prune) or the complete new
meta (meta) in the temp.
The child (host-mode-store-crash-child.ts) intercepted fs through a generic dispatcher: five write paths
(path-based fs.promises.writeFile and appendFile, handle.writeFile,
handle.appendFile, handle.write) rewritten through `as never`/`Function`
casts, plus a rename wrapper. Running the previous child with each path
labelled showed that across the nine windows only handle.writeFile (prune's
`.log.tmp-*` rewrite and meta's `.meta.tmp-*` write) and handle.appendFile
(append's frame) ever fire; the path-based writes and handle.write never do.
armCrash is now two explicit typed wrappers chosen by the op being crashed,
tearWriteFile (handle.writeFile) and tearAppendFile (handle.appendFile), plus
one rename wrapper; the interception targets are narrowed to the exact file of
each op (the temp sibling for prune and meta), the `armed` flag is gone because
the wrappers are only installed after the setup writes, and the types are the
real fs/promises ones with no casts on the wrapped methods (fs.promises is
patched through a member-typed alias, which type-checks without one). The
kill is still child-local, the write is still torn at half its payload, and
the process still dies by a real SIGKILL.
Same nine crash-window cases before and after (9/9 pass). Scratch mutations
that move a crash point each fail the windows they should: before-rename
crashing after the rename (3 cases), after-rename crashing before it (3),
mid-write tearing the whole payload (3) or zero bytes (3), and an interception
aimed at the wrong file or method (the crash point is never reached: 1, 3, 2,
2 cases).
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… fail, share node lifecycle helpers, and refuse to signal a PID that is not a daemon of this checkout
1. The no-reuse check
Devnet suite (swm-host-store-durability): the no-reuse check could not fail on
a recycled seqno. It filtered the recovered log by the pre-kill high-water mark
and then checked that the remaining suffix was sorted. That discards exactly the
frames that would show reuse: with [1,2,3,4] and meta 3 before the kill, a
recovery that dropped complete frame 4 and appended different frames 4 and 5
passed every assertion (sorted, unique, fresh = [5], suffix = [5]) although
seqno 4 was recycled.
The suite now snapshots the complete frames on disk at the kill (seqno plus a
sha256 of each whole frame: timestamp, seqno, length and ciphertext) and the
.meta cursor, and reads them only after the killed processes are really gone
(an unreachable API alone does not say a write already in flight has landed).
After the restart it requires:
(a) the recovered log starts with exactly that preserved prefix: none of the
frames complete at the kill dropped, replaced in place or reordered;
(b) every frame after that prefix has a seqno strictly above the high-water
mark max(meta cursor, last complete frame) at the kill, strictly
increasing among themselves, and at least one exists.
The comparison is a pure function in its own module (log-frames.ts:
parseLog, checkNoSeqnoReuse) with no-devnet tests (log-frames.test.ts, 15
cases, added to this suite's vitest config include like the mixed-version and
v10-rs-wallet-rotation suites do for their pure helpers). They include the
assertion-level regression from the review (before [1,2,3,4] with meta 3; a
recovery that drops 4 and appends different 4 and 5 fails, and the legacy
assertion is shown to accept it), the honest recovery (keeps 4, appends 5)
passing, a frame replaced in place, a dropped prefix frame, a dropped tail, a
reordered prefix, a reused seqno above the mark (also when the cursor was
ahead of the log), a duplicate or out-of-order new seqno, and no new frame.
Every existing assertion (no torn tail, no duplicate seqno, sorted, cursor not
below the log tail, temp files swept, host mode re-engaged) is kept. The check
assumes the retention limits cannot prune during the run (the suite's graph is
far below them). Each of seven scratch mutations of the check (digest not
compared, no prefix, cursor ignored, off-by-one at the mark, duplicates
accepted, no-new-frame allowed, dropped tail unreported) fails at least one of
the 15 tests. README updated.
Live evidence at the final head (one devnet session, 6 nodes, 4 cores):
pnpm test:devnet:swm-host-store-durability passed, 3 live tests plus the 15
no-devnet tests (229 s). Every one of the five kills landed in the append ->
cursor write window (four) or inside a durable meta write (one, which left a temp
file that the restart swept), and every cycle reported "preserved N frames byte
for byte, M new above N". Run earlier at 7d5e55045 with the same check: a dist
whose recovery drops the last complete frame fails the live suite in cycle 1
("frame #4 was seqno 4 at the kill but is seqno 5 after recovery"), a log the old
assertion accepts by construction; the check has not changed since.
2. Node lifecycle helpers and the PID guard
The new durability suite carried its own copy of daemon lifecycle knowledge that
core-peers-features already maintained: the PID-file convention, liveness,
dead-PID cleanup, the port environment `scripts/devnet.sh restart-node` needs,
and restart plus readiness. A change to any of these needed two edits. It now
lives in devnet/_bootstrap/node-lifecycle.ts, consumed by both suites; the
crash timing and every SWM / replication assertion stay in the suites.
The module keeps the ownership rules of both suites: a PID is only signalled if
it was read from a PID file of this devnet's own node home, and a PID file is
only removed when its PID is known dead (a live process keeps its file until it
is explicitly stopped). The kill in the durability suite is still an immediate
SIGKILL of every live daemon process the node's PID files list, and
core-peers-features still escalates SIGTERM to SIGKILL in its own
stopNodeProcesses, which stays in the suite and now sits on the shared
primitives.
Where the suites differ, the module takes the choice as an explicit argument
instead of unifying it:
- Hardhat port for restart-node: `rpcUrl` (core-peers-features passes RPC, which
honours DEVNET_RPC, then node1's chain.rpcUrl, then the default;
swm-host-store-durability passes node1's chain.rpcUrl only, via
rpcUrlFromNode1Config). API and libp2p bases still come from node1's config.
- Readiness probe: NodeProbe { authToken, timeoutMs } (core-peers-features sends
its bearer token with no timeout, polled every 2 s; the durability suite sends
no token with a 3 s timeout, polled every second).
- An unparseable PID file: core-peers-features leaves it, the durability suite
removes it with the dead ones; clearDeadNodePidFiles takes removeUnparseable
(default false) and the durability suite passes true.
Differences that remain and are not byte-for-byte:
- swm-host-store-durability's probe moved from fetch to the plain node:http GET
core-peers-features has always made (same 200 check, same 3 s abort), so the
request now carries a JSON Content-Type header on the GET.
- swm-host-store-durability's node1 rpcUrl fallback moved from `??` to a truthy
check: an empty chain.rpcUrl now falls back to the default instead of throwing
in `new URL('')`.
- core-peers-features drains the status response instead of accumulating and
JSON-parsing it (only the status code was ever used), takes waitFor from the
harness inside the shared wait helpers (same loop and error text as its local
copy, which it keeps for its own waits), and reads node1's config for the RPC
through the shared reader with the same precedence (DEVNET_RPC, config,
default). Its request, headers, intervals and timeouts are unchanged; its kill
of the victim core now goes through the verified kill below.
A live PID is signalled only if it is a daemon of this checkout. Both previous
copies signalled every live PID read from the PID files, so a stale file whose
number an unrelated process had taken would have had that process SIGKILLed. A
live PID is now signalled only if its command line (`ps -ww -o command=`) is a
DKG daemon started from this checkout: the repo root as a path (as given, or as
its realpath) and a last argument of daemon-supervisor, daemon-worker or
daemon-foreground-worker. A live PID that fails the check, or whose command line
cannot be read, makes verifiedNodePids / sigkillNodeProcesses throw, naming the
file and the PID and never the command line, before anything is signalled: a
silent skip would leave the node running and turn the kill -9 into a graceful
stop. A dead PID is skipped as before.
- Why the checkout and not the node's home: observed with a real
daemon-supervisor and its daemon-worker, argv is
`node <repoRoot>/packages/cli/dist/cli.js daemon-supervisor|daemon-worker`;
the home (DKG_HOME) is only in the process environment. devnet.sh starts nodes
with DKG_NO_BLUE_GREEN=1, so the entry point is always the checkout's own CLI
(mixed-version nodes live under <repoRoot>/.devnet-versions). A check on the
home would need `ps`'s environment output and the exact text devnet.sh
exported, and a refused live daemon would silently turn a kill -9 into a
graceful stop.
- What it cannot tell apart: a stale PID recycled by another daemon of this same
checkout, for example another node of the same devnet. The PID files are then
the only per-node distinction.
- The durability suite verifies the host's PIDs before its wait for the next
frame and passes them to sigkillNodeProcesses({ verified }), so the kill at the
frame stays immediate (the check costs a `ps` per process, tens of
milliseconds, and the kill window is a few milliseconds wide).
core-peers-features' stop and victim-kill paths go through the same check.
- A stale claim copied from the base is corrected: devnet.pid is not the
`cli.js start` launcher that exits after forking. devnet.sh's
refresh_node_pidfile writes the detached daemon-supervisor (the worker's
parent) there, or the worker itself when it has been reparented to init.
New no-devnet tests (devnet/_bootstrap/node-lifecycle.test.ts, 40 cases, added
to the include of vitest.manifest.config.ts next to harness-json.test.ts, so
`pnpm test:devnet:manifest` runs them): PID parsing and de-duplication, the
ownership rule with real child processes (a bystander and another node's process
are never signalled, a dead PID in a file is not, a live process keeps its file),
per-node ownership (node homes 1 to 5 exist and node1 holds live decoy PID files:
reading, clearing and killing node2 and node3 see only their own files), the
daemon check (a table of accepted and rejected command lines, the recycled-PID
case with an injected reader and with the real `ps`, both spellings of the root,
the refusal of the verified worker when the other file lists a foreign PID,
no command line in the error), the verified-PID shortcut for time-critical
kills, the two port policies side by side, the readiness probe against a local
server (token sent only when asked, non-200 and refused are false, a hung node
gives up once a timeout is set) and restartNodeAndWait against a stand-in
devnet.sh (args, cwd, port environment, polling, the timeout error, no waiting
when the script fails). Twenty-four scratch mutations of the module, re-run at this head, each fail at
least one of them: among them a nodeHome that resolves every node but 4 to node1
(which survived the first 20 tests), the node number ignored, no command-line
check, any root accepted, the subcommand not required or not last, no path
boundary, an unreadable command line skipped, the verified list ignored, no
realpath spelling, PIDs not de-duplicated, dead PIDs signalled or looked up,
live-PID files deleted, a one-file read, 4xx accepted, the token or timeout
ignored, and the port environment dropped.
Live evidence at the final head (one devnet session): all twelve PID files of the
real 6-node devnet verify (devnet.pid is the daemon-supervisor, reparented to
init; daemon.pid is the daemon-worker whose parent is that supervisor);
core-peers-features passed 6/6 (102 s, core-fill in 42 s) with its stop and
victim-kill paths on the verified kill; the durability suite's five kill cycles
passed with the verified PIDs. Earlier, at 7d5e55045 (the extraction before the
guard), core-peers-features passed 6/6 in both the previous version and the
refactored one, each twice in one session (85 s and 50 s before, 59 s and 86 s
after; no test differed in outcome). The known core-fill flake did not occur in
any run, so nothing is known about its rate here.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…at the review found The header of devnet/_bootstrap/node-lifecycle.ts now says that devnet.sh honours a DEVNET_VERSIONS_DIR override (with one outside the checkout a healthy mixed-version daemon would be refused, loudly; none of the suites here run mixed-version nodes), and that PIDs are verified once and signalled later with only a liveness re-check at the signal, so a verified PID that exits and is recycled inside that window of seconds would still be signalled. Comment text only. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…t-store-durable-writes # Conflicts: # pnpm-workspace.yaml
POST /api/shared-memory/host-catchup is the operator's way to confirm what a hosting core holds for a curated graph. It reported counts and a cursor only, and always asked for the host's default page. Two optional body fields: - maxEntriesPerRound asks each host for smaller pages (the agent and the wire request already supported it); - includeEntries adds, per peer, the seqno and SHA-256 of every envelope that peer served, in order, whether or not the replay applied it. Without them the request and the response are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… in host catch-up Three gaps in the swm-host-store-durability suite's last scenario: - The kill cycles allow the .meta cursor to trail the log by the frame being appended, which also accepts a cursor that is never persisted. With ingestion stopped the cursor must now cover the whole log (checkCursorCoversLog). - maxRounds limits requests per call, not entries per page, so paging could finish in one data response. The scenario now asks for a page a third of the log and requires at least three non-empty pages from 0. - A matching count and final cursor do not show which frames arrived: [1, 1, 3, 4] for a log of [1, 2, 3, 4] has both. The served envelopes are now compared, in order, with the frames on disk after the cursor, seqno and ciphertext digest (checkServedFrames). Both checks are pure functions with no-devnet tests for the cases above. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
POST /api/shared-memory/host-catchup passes maxEntriesPerRound through to catchupSwmFromHost. The request digest binds maxEntries, and the requester signed the value it was asked for while the encoder clamps it to 1024, so a page size above 1024 produced a signature the host could not verify and the catch-up fetched nothing. normalizeCatchupMaxEntries applies the encoder's default and clamp before signing, so the signed and the sent value are the same. Wire tests: a request signed for 2048 and sent as 1024 fails to verify; the normalized request verifies at the limit and above it. The live suite's quiescence gate passed without waiting: its first probe compared the log with a read taken just before it. quietPeriodGate now requires the same observation for LOG_QUIET_MS (6 s) and restarts on any change. It observes the log size together with the .meta cursor, since a frame is durable before its cursor: a devnet run under load read cursor 2 with three frames on disk in the suite's first test, which now waits on the same gate before it compares them. Live run on a six-node devnet: 33/33 (30 unit tests of log-frames, the baseline, five kill -9 cycles, paging from four cursors). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sends The page-size tests in host-catchup-wire.test.ts normalize the page size themselves, so they stayed green with the normalization removed from catchupSwmFromHost. The new file runs the real method with a real signer and a stand-in messenger, decodes what was sent and verifies it as the host does: for a page above the wire limit, at it, a small one, a fractional one and the default. With the previous assignment restored in catchupSwmFromHost, the oversized case fails with "signer mismatch". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The SWM host store built a context graph's metadata from defaults whenever the .meta read failed, whatever the error: EIO, EACCES and EMFILE counted as a missing file. The log-tail recovery did the same for the .log, through an existence probe and a read that both answered "no log" on any error. The result was cached, and the next mutation wrote it back over the file. Since a failed .meta write drops the cache, a warm process can go through that load again, so one failed write followed by one transient read error was enough. With the log pruned to empty the cursor restarted at 0: the next append was acknowledged as seqno 1 and the rewritten .meta no longer carried hostModeSubscribed, so a requester resuming from the old cursor never saw the new frames and a restart did not re-engage the subscription. With the cursor behind the log (the frame of a failed cursor write), a failed log read let the lagging cursor stand and the next append wrote a seqno that was already in the log. And a registered context graph loaded as unregistered, so the next prune applied the unregistered TTL and byte cap to it. A load now takes only a missing file (ENOENT) as absent: no .meta means defaults, no .log means no frames. Any other read error on either file rejects the load and the mutation that needed it; nothing is cached, nothing is persisted, and the next access reads the files again. The log tail is recovered by the read alone, without the separate existence probe. The cache eviction and the failed-mutation rollback are unchanged. So is a .meta that reads but does not parse: it stays unusable (defaults on load, reaped with its log by init()). The unlocked readers still swallow the rejection and answer that one call as if the context graph were unknown. The on-disk format, the wire format and the method signatures are untouched. Tests, in host-mode-store-durability.test.ts, eight cases: after a prune-to-empty and a failed meta write, EIO, EACCES or EMFILE on the re-read (one case each) rejects the append, writes nothing, and the next append is seqno 6 with the flag intact across a restart; a .meta that is really missing on the same route still appends from seqno 1; a transient error is not cached, the next call on the same store reads again; a reader that hits the error leaves no defaults behind for the next prune; a failed log read rejects instead of writing a seqno twice; a failing existence probe cannot hide the log. With the previous source seven of the eight fail and the ENOENT control passes: the scenario with 'promise resolved "1" instead of rejecting', the log-tail case with 'promise resolved "2" instead of rejecting', the prune case with 63 bytes pruned where 0 were expected. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…h it cannot load With a load that rejects on an unreadable .meta or .log, prune() stopped at the first context graph whose load failed and left every one after it unpruned. A context graph that stays unreadable (a damaged .log next to a readable .meta) would then keep the later ones from ever being pruned. prune() now skips a context graph whose prune rejects, goes on with the rest, and rethrows the first failure once the sweep is done, so the prune timer still logs it. Test: two context graphs with every frame expired, EIO on the first log read. prune() rejects with EIO, the other context graph is pruned in that same sweep, and the next sweep prunes the one that was skipped. Before this change neither was pruned. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t-store-durable-writes # Conflicts: # devnet/_bootstrap/vitest.manifest.config.ts # devnet/suites.json # package.json # pnpm-lock.yaml
…le-size ratchet passes testnet-canary now fails CI when a baselined file grows past its budget (node scripts/audit-file-size.mjs). The host catch-up page size and includeEntries fields (f4d77b1) took packages/cli/src/daemon/routes/memory.ts from 2,102 to 2,115 lines against a budget of 2,102. The whole POST /api/shared-memory/host-catchup handler now lives in packages/cli/src/daemon/routes/shared-memory-host-catchup.ts (80 lines) as handleSharedMemoryHostCatchupRoute(ctx), which returns false for requests it does not own; memory.ts dispatches to it where the block stood. The body is the same: parsing, defaults, the 501 and 500 responses and the response shape are unchanged (each `return jsonResponse(...)` became `jsonResponse(...); return true`). memory.ts is 2,055 lines (2,115 before, 2,102 at the base), and its budget in scripts/file-size-baseline.json is lowered from 2,102 to 2,055 (only that entry: node scripts/audit-file-size.mjs --write would also have lowered 50 unrelated entries that other merges had already shrunk). Measured: node scripts/audit-file-size.mjs passes for this file; cli test/shared-memory-catchup-durable.test.ts (33 tests, which drive the route through handleMemoryRoutes) and node-ui test/openclaw-bridge.test.ts + test/ui-compat.test.ts (which read the daemon source) pass; cli tsc --noEmit is clean. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…mode-store.ts 1,164 -> 615 lines) testnet-canary's file-size ratchet (node scripts/audit-file-size.mjs) refuses a new source file over 800 lines, and the host store had reached 1,164 after the load-error rule and the prune sweep fix. It is split by concern; the code moved, it was not rewritten: - host-store-durable-fs.ts (242 lines): DurableFiles, the crash-safe file operations (writeFileDurable, unlinkDurable, appendFileDurable, truncateFileDurable), the pending directory-fsync bookkeeping (pendingDirSync, per-change generations, completePendingDirSync) and the temp file lifecycle. Ported from the earlier unpushed round (3facd60fa) onto the current store. - host-store-format.ts (158): file naming (HostStoreLayout) and the frame codec (encodeFrame, scanLogFrames, readEntriesSince, countFrames, planRetention, concatFrames). Pure functions; every reader keeps the break-on-overrun walk. - host-store-meta-loader.ts (152): the metadata cache and the one cold-load owner per CG (load, initialize, recoverLastSeqnoFromLog), with the ENOENT-only rule for a missing file and the no-await registration step kept as it was. SwmHostModeStore.loadMeta stays as a non-async delegate returning the loader's promise. - host-store-startup-reconcile.ts (143): reconcileOrphanLogs(dataDir, files), the init sweep of orphan logs, corrupt metas and stale temps. - host-store-types.ts (124): the public types, the default limits and CgMetaState. host-mode-store.ts re-exports the public types and META_FILE, so no importer changes. - host-mode-store.ts (615): the public API, the per-CG write lock, sequence allocation, the .meta flags, persistMeta, the tail-repair decision and the retention policy. Its header lists the modules. The other agent's changes are intact: a load rejects on any read error but a missing file (1aec765) and prune() keeps sweeping past a CG it cannot load (695e68e). Order is untouched: the frame is fsynced before persistMeta, a failed cursor write drops the cache, an idempotent retry still calls completePendingDirSync. Behavior preservation, measured. A scratch differential harness (not committed) drives the merged head's host-mode-store.ts (1,164 lines, copied) and this tree over seeded op sequences on an in-memory filesystem that records every fs call (order and arguments), every file-handle call and every directory fsync, and completes each call on its own macrotask in a seeded order so concurrent ops interleave. Sequences mix append, markRegistered/Unregistered, markHostModeSubscribed/Unsubscribed, prune (with a TTL and a byte cap small enough to trim and unlink), iterate, stats, the unlocked readers, init/reconcileOrphanLogsNow, clock jumps and restarts with seeded crash damage (torn tails, lagging or corrupt metas, orphan logs, stale temps, decoys). 42,388 paired runs (sequential and concurrent; random failures at 2 to 50% of calls; a single failure injected at every call index of 200 sequences x 2 modes): 647,492 ops, 4,404,439 fs calls per implementation, 214,698 injected failures. Every trace (calls, completions, results, thrown error codes and messages, startup reports) and every final file set was identical. The unchanged host-mode-store, host-mode-store-durability (73 tests), dirsync-model and key-canonicalization suites pass (112 tests, no test file edited for this commit); agent tsc --noEmit is clean. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…wnership before restarts, and unit-test the cursor and paging accounting Three review findings on the devnet lifecycle helpers, and two gaps found by comparing the earlier session's unpushed suite changes with what landed. Ownership check (devnet/_bootstrap/node-lifecycle.ts). isDaemonOfCheckout accepted any process whose command line held a path under the checkout and whose last argument was named like a daemon command, so `node <checkout>/node_modules/vitest/vitest.mjs run daemon-worker` passed and would have been SIGKILLed from a stale PID file. It now requires the checkout's CLI entry point (`<root>/packages/cli/dist/cli.js`, as its own argument, under either spelling of the root) directly followed by daemon-supervisor, daemon-worker or daemon-foreground-worker as the last argument, which is what resolveDaemonNodeCommand starts (`<node> [execArgv] <entry> <subcommand>`). The entry point is searched for in the `ps` text rather than in whitespace-split tokens, so a checkout path with spaces still matches. A flag that carries the entry path (`--import=<entry>`) no longer qualifies; that accept case of the old test is gone. Restarts (node-lifecycle.ts, swm-host-store-durability, core-peers-features). `devnet.sh restart-node` stops the node by signalling whatever its PID files list, with no check, and the suite's setup reached it for node4 and node5 after editing their configs. restartNodeAndWait now runs the new stopNodeProcesses first: verifiedNodePids (throws before anything is signalled if a live PID-file entry is not a daemon of this checkout), SIGTERM, SIGKILL after the grace period, then the dead PID files are removed, so the script's own stop phase finds no live entry to signal. core-peers-features' private copy of that stop is replaced by the shared one (same 10 s + 10 s windows). The swm-host-store-durability setup also verifies both nodes before it edits any config. Not changed: the script also sweeps the process table for processes that mention the node's home directory (its managed store servers and detached children); that sweep is not PID-file based and is documented in the module header and the suite README. Node-based API (node-lifecycle.ts). The optional `verified` mode of sigkillNodeProcesses turned a node-based call into a raw PID-list call; it is gone. sigkillPids returns the PIDs it signalled, the suite verifies with verifiedNodePids before its wait and calls sigkillPids at the kill point, and sigkillNodeProcesses is verify-then-sigkillPids. Suite assertions (log-frames.ts, automated.test.ts). Ported from the unpushed round, as the parts the landed versions lack: - checkCursorCoversLog takes `appendInFlight`: while the writer runs, the cursor may trail by the one append in flight, counted in frames. The kill cycles used `cursor >= last seqno - 1`, which is measured in seqnos: with a seqno burned by a failed append (log 7, 8, 10, cursor 8) it rejects a cursor that is exactly one frame behind. The quiescent form is unchanged. - checkCatchupWalk accounts for a walk page by page (each call resumes at the previous cursor, a page holds at most the page size and advances the cursor over exactly the frames it served, an empty page ends the walk, exactly ceil(frames / page) nonempty pages, and at least 3 from 0), so the claim that paging crosses page boundaries is unit-tested instead of living only in the live loop; one full response followed by an empty one fails it. Not ported: the earlier round's per-cycle quiescent re-check (it stops and restarts the writer in every cycle for a state the final quiescent check already covers) and its large-share page planner (the landed suite pages with maxEntriesPerRound). Tests, measured: node-lifecycle.test.ts 52 tests pass (was 40), including the vitest example and `ps`-based cases with real child processes (an unrelated live PID in devnet.pid rejects restartNodeAndWait before devnet.sh runs and leaves it and the legitimate worker alive; the verified stop signals the worker and supervisor and leaves bystanders and other nodes alone; SIGKILL escalation for a process that ignores SIGTERM); log-frames.test.ts 39 tests pass (was 30). A throwaway tsc over the changed devnet files reports only the pre-existing devnet/_bootstrap/harness.ts HDNodeWallet error. This commit was not run against a devnet. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…o four focused files with shared support
Review thread: host-mode-store-durability.test.ts held seven concerns under one
fixture and rebuilt its pause machinery inline. It is now four suites along its
existing describe boundaries, plus a support module (ported from the earlier
unpushed round onto the current file, which the other agent had extended with
the "a file the store could not read is not an absent file" tests):
- host-mode-store-durable-writes.test.ts (324 lines, 14 tests): the atomic
write sequence and how each step's failure is handled
- host-mode-store-tail-recovery.test.ts (234 lines, 10 tests): leftover temp
files, torn log tail
- host-mode-store-cold-init.test.ts (547 lines, 26 tests): the single-owner cold
load, seqno monotonicity across crash windows, and the unreadable-file
rule (the nine tests the other agent added, moved here)
- host-mode-store-dirsync-retries.test.ts (448 lines, 23 tests): a retry after
a failed directory fsync, including the unlink cases
- test/_helpers/host-mode-store-durability-support.ts (231 lines):
useDurabilityFixture (fresh data dir per test, FileHandle prototype, real
directory fsync, traceDurability, tempFiles, seedLaggingMeta), the frame
helpers, and the gates createGate, gateRenames ({ every } pauses every
rename, default only the first) and holdNextDirSync, which replace
gateFirstRename and the inline copies of it and of holdNextDirSync.
Every suite keeps the outer describe('SwmHostModeStore durable writes') and
declares its own vi.mock of fsyncRfc64DirectoryV1, so each full test name is
unchanged. The four files replace the old one in vitest.unit.config.ts; the
seeded model test's comment points at the retries file.
Preservation, measured: the vitest JSON reporter over the old file (at the
merged head) and over the four new files lists the same 73 full test names
(sorted lists equal), all passing. The 287 lines that carry an assertion are
identical across old and new modulo the `fx.` prefix, except the one line that
races the gate's `reached` promise.
Mutation checks, measured on the extracted modules (src/swm/host-store-*.ts,
host-mode-store.ts), each applied alone to the committed code and run against
the new suites plus the store, model and key-canonicalization suites, against
the old file, and against the differential harness (28 mutations):
- the four mandated ones: no generation check (3 tests: 2 durability, the seeded
model), no completePendingDirSync (DurableFiles: 19; store call sites: marks 15,
prune with no log 3, prune with nothing to drop 2), the cursor written before
the frame (7), the load-error rejection dropped (meta read 5, log read 2);
- the rest of the earlier round's list: no directory fsync after an unlink (7) or a
rename (29), no temp cleanup (4), no live-temp guard (1), no pending mark (19),
no temp fsync (6), no frame fsync (2), no tail truncate (4), directory snapshot
taken after the fsync (5), unshared cold initialization (12), a rejected
initialization cached (13), the cache installed before the reconcile write (7),
the reconcile write not best-effort (3), no log tail in the cursor (21), no
tail re-check after a failed append (1), no seqno reservation (2), the cache
kept on a failed cursor write (8), no tail repair (4), the prune sweep stopping
at the first failure (1).
Each fails tests, and the old file and the new files fail the same durability
tests every time. One mutation passed every suite: init()'s sweep reaping a
.log whose .meta it could not read (EMFILE/EACCES), a rule from an earlier
review round that no test pinned; host-mode-store.test.ts now has a test for
it, which fails that mutation (1 test). The differential harness (scratch)
detects 27 of the 28 at its default settings; the one it misses is the missing
generation check, which only the two deterministic overlap tests and the
model catch.
Agent unit lane: the four files (73 tests), host-mode-store (27 tests, one new),
dirsync-model and key-canonicalization pass.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Moving the route into its own module (shared-memory-host-catchup.ts) turned all of its lines into changed lines for the changed-line coverage gate, and a partial run (the 33 existing route tests) covered 21 of 26 changed executable lines, 80.8% against the 80% gate for the cli package. The five uncovered lines were the error answers, which nothing tested: - a missing, blank or non-string contextGraphId answers 400 and never reaches the agent (4 cases); - an agent build without catchupSwmFromConnectedHosts answers 501; - a catch-up that throws answers 500 with the agent's message. shared-memory-catchup-durable.test.ts has 39 tests (was 33) and passes. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… a restart right after a SIGKILL is not refused The first live session of the previous commit (6 nodes, real daemons) showed a regression in it. A preflight over all six nodes passed, and the new suite's three live tests and the 39 no-devnet tests passed, but its afterAll could not restore node4 and node5: cleanup: could not restore node4: node4: daemon.pid lists pid 64337, which is alive but is not a DKG daemon started from <checkout> ... and core-peers-features, which runs next on the same devnet, then failed at setup with ECONNREFUSED 127.0.0.1:9204 (node4 was down). The cleanup does sigkillNode(num), clearDeadNodePidFiles(num), restartNodeAndWait(num). The SIGKILLed worker is a zombie until its parent (itself killed) or init collects it: it still answers signal 0, so pidAlive said alive and the PID file stayed, and `ps` shows it as `(node)`, so restartNodeAndWait's new verified stop refused it as "not a daemon of this checkout". Before that stop was added, the shell's restart-node tolerated it. The kill cycles never saw this because they wait for the killed PIDs to go before restarting. pidAlive now reads the process state (`ps -o stat=`, new readProcessState) and counts a zombie as gone; a state that cannot be read counts as alive. That also makes waitForPidsGone and clearDeadNodePidFiles treat a zombie as gone, and verifiedNodePids re-checks liveness before it refuses a command line that is not a daemon's, so a process that exits while it is being looked at is skipped. Tests (node-lifecycle.test.ts, 57 tests, was 52), each with a real zombie (a `sleep 0.1 &` whose parent, an `exec sleep`, never waits for it): pidAlive and readProcessState on a zombie, clearDeadNodePidFiles removing its file, verifiedNodePids and sigkillNodeProcesses skipping it, stopNodeProcesses next to a live supervisor, and restartNodeAndWait with a zombie in daemon.pid reaching devnet.sh. A mutation that makes pidAlive ignore the state fails 5 of them. pnpm test:devnet:manifest: 143 tests pass (5 files). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…e, not only a zombie The previous commit counted a zombie (`Z` state) as gone, on the diagnosis that a SIGKILLed daemon is still a zombie when swm-host-store-durability's afterAll restarts it. A second live session (6 nodes) failed the same way: node4 and node5 were not restored, and core-peers-features then failed at setup with ECONNREFUSED 127.0.0.1:9204. A temporary debug line in verifiedNodePids (not committed) recorded what it refused: the daemon.pid worker of node4 and of node5 with the command line `(node)` and the `ps` state `?E`, raised from restartNodeAndWait's stop right after the cleanup's SIGKILL. On macOS the `E` flag means "the process is trying to exit": it is in the middle of exiting, not yet a zombie, still answers signal 0 and shows `(node)` instead of its argv. processStateIsGone(state) is now true for a `Z...` state or one with the `E` flag, and pidAlive uses it (ordinary states such as `S`, `Ss+`, `R+` and an unreadable state still count as alive). node-lifecycle.test.ts has a unit test of the state check with the strings seen on the live devnet, and the real-zombie tests of the previous commit still apply; 58 tests pass, pnpm test:devnet:manifest passes. An exiting process cannot be held in that state on demand, so the live devnet session is what exercises it. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…aged store processes `devnet.sh restart-node` stops a node by scanning `ps eww` for anything that mentions the node's home directory and signalling every hit (SIGTERM, then SIGKILL after 15 s) plus its children. A `tail -f <home>/daemon.log` (what `devnet.sh logs <n>` runs), an editor or a grep that only names the home was signalled too, so a restart could kill a developer's or another agent's bystander process on a shared machine, after the TypeScript side had verified the PID-file entries. The sweep now keeps a hit only when its executable is `node` or a managed store binary (`oxigraph*`). PID-file entries and children of kept processes are unchanged. Tests run the real script: collect_devnet_node_pids lists a node process and an `oxigraph-v*` process that mention the home and not a log tail or another node's process; stop_devnet_node_processes reaps the first two and leaves the tail; and restartNodeAndWait through the real restart-node (start_node replaced by a recorder) leaves the tail alive. With the filter reverted three of the four fail. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
| * process-table sweep is limited to node and managed store processes. | ||
| */ | ||
| export async function restartNodeAndWait(paths: DevnetPaths, options: RestartNodeOptions): Promise<void> { | ||
| await stopNodeProcesses(paths, options.num, { readCommandLine: options.readCommandLine }); |
There was a problem hiding this comment.
🟡 Issue: Give node stopping one canonical implementation
What's wrong
This adds a second shutdown implementation in front of the existing shell shutdown rather than extending the canonical lifecycle boundary. The two implementations have different process sets, liveness rules, and grace periods, and callers must understand their ordering to understand a restart. The existing restore helper adds another redundant stop.
Example
Restoring an unreachable core now traverses three stop phases: restoreNodeIfDown's stop, restartNodeAndWait's stop, and devnet.sh's stop. The TypeScript and shell implementations separately own discovery, signalling, escalation, waiting, and PID-file cleanup.
Suggested direction
Put ownership verification and shutdown behind one process-control implementation used by both suite helpers and the script. Restart should perform one checked stop, start the node, and probe readiness.
For Agents
Consolidate process control across node-lifecycle.ts and scripts/devnet.sh, then remove redundant stop orchestration from core-peers-features. Preserve ownership verification before signalling, daemon and managed-store cleanup, escalation, and readiness probing. Use the existing lifecycle tests to verify those behaviors.
|
|
||
| export class HostStoreMetaLoader { | ||
| /** The metadata installed so far, by `cgKey`. The store sets and drops entries as it persists. */ | ||
| readonly cache = new Map<string, CgMetaState>(); |
There was a problem hiding this comment.
🟡 Issue: Keep metadata cache and persistence under one owner
What's wrong
The extraction splits a tightly coupled metadata lifecycle between two classes while exposing its backing Map as their coordination mechanism. The loader owns initialization and reconciliation writes, but the store owns cache replacement, eviction, and authoritative writes. This moves code without establishing a clean ownership boundary, leaving future metadata changes dependent on both classes' internal bookkeeping.
Example
After a metadata write fails, SwmHostModeStore deletes the loader's cache entry. The next load runs reconciliation and a best-effort write inside HostStoreMetaLoader, then a subsequent mutation writes through SwmHostModeStore again. Understanding one metadata retry requires following persistence and cache transitions across both classes.
Suggested direction
Make the cache private and move authoritative metadata persistence and its failure eviction into the metadata component. Expose operations that maintain those invariants instead of exposing the Map.
For Agents
Make the metadata component own loading, persistence, cache installation, and failure invalidation; keep sequence allocation and retention policy in SwmHostModeStore. Preserve burned sequence numbers, shared cold initialization, best-effort reconciliation, and directory-sync retry behavior. Run the existing cold-init and durability retry suites.
…ntains a space
The sweep that limits restart-node's process-table scan to node and managed
store processes judged a process by the first word of its command line. For an
executable under a directory with a space in its name ("My Projects") that word
is a path prefix, so the node's own processes and its oxigraph store would not
have been reaped and the restart would have found the data directory still held.
When the first word does not name a node or oxigraph binary, the executable that
`ps -p <pid> -o comm=` reports decides. A test runs a node whose path has a space
in it; without the fallback it fails.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
| expect(log.seqnos[0]).toBe(1); | ||
| // The API view agrees with the disk (cleartext or wire-id key). | ||
| const stats = await hostStats(); | ||
| const entries = Object.values(stats?.perCg ?? {}).reduce((sum, row) => sum + row.entries, 0); |
There was a problem hiding this comment.
🔴 Bug: The API consistency assertion can pass using unrelated graphs
What's wrong
The comment claims that the API agrees with the tested graph's disk contents, but the assertion checks only a store-wide minimum. The disk gate remains meaningful; this API check can conceal a missing graph or an incorrect count.
Example
The tested graph has three frames on disk, but the API returns perCg: { unrelated: { entries: 3 } }. This assertion still passes although the tested graph is missing from the API response.
Suggested direction
Select the tested graph's perCg row, require it to exist, and compare its entry count with log.frames.length.
For Agents
Update the baseline stats check in automated.test.ts to resolve the tested graph's cleartext or wire ID and compare its row with the parsed log. Add a negative case where unrelated graphs have enough entries but the tested graph is absent.
Summary
SwmHostModeStore(the per-CG opaque-ciphertext log plus.metacursor that hosting cores keep under<home>/swm-host/) rewrote its files with a plainfs.writeFileand never fsynced anything. Akill -9or power loss inside the prune rewrite or a meta write could leave a truncated log or a torn/empty meta, and the "crash betweenappendFile(durable) andpersistMeta(durable)" reasoning inloadMetarested on writes that were not durable. The prune failure reproduces on the base commit: a child process SIGKILLs itself mid-rewrite and the log is left holding a truncated prefix of the pruned bytes (seqnos 5-6 of 5-9), with the acknowledged entries gone.User-facing impact. Operators hosting SWM context graphs on cores: a power loss or
kill -9during a prune, a meta write or an append no longer risks a truncated log, a lost cursor, later entries buried behind a torn frame, or a recycled seqno (which breaks host catch-up's strict-greater-than paging), and a log that prune removed because every entry expired stays removed. Cost: each acknowledged append now does three fsyncs (see the decision below), and a prune that removes a whole log does one directory fsync.Changes
persistMeta, theloadMetareconcile write and the prune log rewrite now go through onewriteFileDurable: sibling temp file, fsync the temp handle,renameover the target, thenfsyncRfc64DirectoryV1on the directory; the temp file is removed on failure.persistMetastays authoritative (rejects), the reconcile write stays best-effort, and the all-dropped prune branch removes the log and then directory-syncs the unlink (below).<key>.<log|meta>.tmp-<pid>-<uuid>. The existing scans key off the.log/.metasuffix, so they already ignore it, andinit()now deletes them: never a temp this same instance still has in flight, never counted as an orphan log, and it does not fire the operator warning on its own (reported as the optionalstaleTempFilesRemoved).ENOSPCmid-append) made every later append invisible toiterate()and let seqnos be reused. Reproduced on the base commit: after a partial frame, two appends were acknowledged as seqno 3 and 4 anditerate()still returned only 1 and 2. The tail is now checked once per CG per process, under the per-CG write lock, and truncated to the last complete frame; it is re-checked after any failed append. Readers already stop at that boundary, so no servable byte is lost. A failed append now burns its seqno instead of risking a duplicate frame.persistMetadrops the cached meta when its write fails (a retriedmarkRegisteredused to no-op against a flag the disk never saw).loadMetahas one owner per CG: the first caller starts a cold initialization (read.meta, recover the log tail, take the max of the two cursors, best-effort persist the reconciled cursor, install into the cache) and every other caller, unlocked readers and locked mutators alike, awaits that same promise, so no mutation can interleave with it. The earlier cache-install guard only kept a slow cold load from rolling back the in-memory cursor; it did not stop the delayed reconcile rename from overwriting a newer.metaon disk (reproduced: a clearedhostModeSubscribedcame back after a restart, and a fresh store re-engaged the graph). The initialization never takes the per-CG write lock (mutators hold it while they await the initialization), a rejected initialization is not cached, persisting the reconciled cursor stays best-effort, andSwmHostModeStore.loadMetais a non-async delegate returningHostStoreMetaLoader.load's promise (async, body withoutawait), so a non-string id is still a rejection the readers swallow.pendingDirSync, aMap<target, generation>), and an idempotent retry completes that fsync before it acknowledges:markRegistered/markUnregistered/markHostModeSubscribed/markHostModeUnsubscribedwhose flag already matches, and a prune that finds the log already pruned or already gone. A directory fsync covers the changes that had returned before it started, so a fsync snapshots the pending(target, generation)pairs first and, once it succeeded, forgets a target only if its entry still has the snapshotted generation: a target that is changed again and fails its own fsync stays pending while an older fsync is in flight. A failing retry keeps the mark, and with nothing pending no extra fsync runs.pruneremoved a log whose every entry had expired with a barefs.rm, so a power loss could bring expired ciphertext back. The unlink now goes through the same pending-fsync machinery (one directory fsync per removed log, at prune time only). A cap eviction insideappendcan take this branch too (a single frame larger than the per-CG byte cap): if the directory fsync after that unlink fails, the append rejects although its frame and cursor were already durable, and the retry stores the next seqno, the same failure mode as a failing rewrite in the eviction path.node scripts/audit-file-size.mjs).host-mode-store.tsgoes from 1,164 to 615 lines; new inpackages/agent/src/swm:host-store-durable-fs.ts(242:DurableFiles, atomic replace, durable append/truncate/unlink, pending directory-fsync generations, the live-temp set),host-store-format.ts(158: file naming and frame codec, pure),host-store-meta-loader.ts(152: metadata cache and the one cold-load owner),host-store-startup-reconcile.ts(143: the init sweep),host-store-types.ts(124). The code moved, it was not rewritten: the public types andMETA_FILEare re-exported fromhost-mode-store.ts, so no importer changes. cliroutes/memory.tsgoes from 2,115 to 2,055 lines (base 2,102; its budget inscripts/file-size-baseline.jsonis lowered to 2,055) by moving the host catch-up route intoroutes/shared-memory-host-catchup.ts(80). Behavior preservation: a scratch differential harness (not committed) drove the merged head's store and the split one over 42,388 paired runs on a recording in-memory filesystem (647,492 ops, 4,404,439 fs calls per implementation, 214,698 injected failures, a single failure at every call index of 200 sequences), with identical traces, results, errors and final files.devnet/swm-host-store-durability, registered insuites.json,pnpm-workspace.yaml, the rootpackage.jsonand the lockfile importer), and shared devnet lifecycle helpers. The suite's no-reuse check compares the recovered log with a byte-level snapshot of the frames that were complete at the kill (log-frames.ts, 39 no-devnet tests: no-reuse, cursor coverage with the one-append allowance while the writer runs, served frames, the page-by-page catch-up walk, the quiet-period gate).devnet/_bootstrap/node-lifecycle.ts(63 no-devnet tests;pnpm test:devnet:manifestruns 149) is shared withcore-peers-features, which this PR therefore touches: PID-file discovery, liveness, dead-PID cleanup, the kill and restart helpers. Every kill and stop these two suites issue signals only PIDs read from this devnet's ownnode<N>/{daemon,devnet}.pidfiles, and only a live PID whose command line is the checkout's CLI entry point (<repo>/packages/cli/dist/cli.js) directly followed by a daemon subcommand as the last argument (a process that merely runs from the checkout, such as the test runner, is refused); a live PID that is not one (a stale file whose number was recycled) makes the helper throw instead of being signalled, and a process that is exiting or is a zombie counts as gone. Restarts go through the same check:restartNodeAndWaitstops the node itself (verify every live PID-file entry, SIGTERM, SIGKILL after the grace period, remove dead PID files) beforedevnet.sh restart-node, whose own stop phase signals PID-file entries unchecked and so finds nothing live. That stop phase also sweeps the process table for processes that mention the node's home directory (its detached children and managed store servers);devnet.shnow limits that sweep tonodeandoxigraph*processes, so atail -f <home>/daemon.log(whatdevnet.sh logs <n>runs), an editor or a grep that only names the home is never signalled. The check reads the checkout path from the command line, not the node's home, because the home (DKG_HOME) is only in the process environment; a recycled PID that became another daemon of the same checkout is not distinguished.No wire, frame-format, catch-up paging or public-API change. The
max(metaSeqno, lastLogSeqno)cold-load reconcile is unchanged. The RFC-64 owner-only file policy is not adopted (atomicity and fsync only); the temp keeps the default modewriteFileused.Decision: fsync every append (option a)
appendfsyncs the frame (appendFileDurable) beforepersistMetapublishes its seqno, so an acknowledged append is durable in both files.persistMetais now durable and runs at the end of every append, the seqno cannot be reused after a crash under either choice (a durable cursor ahead of a lost frame just leaves a gap, which strict-greater-than paging tolerates; pinned by a test). What (c) would lose is the frame itself: a hosting core could acknowledge an envelope that a power cut then drops, which is the durability gap this change exists to close, for one more fsync on top of the two thatpersistMetaalready costs.Measured, single CG, 500 sequential appends of 4 KiB, on macOS/APFS where Node's
fsyncis a full flush (an upper bound relative to a Linux SSD):The two meta fsyncs the task mandates dominate; (a) adds about 75% on top of (c) on this machine. Appends are already serialized per CG behind the write lock, so one CG's sustained ingest is bounded at roughly 60 per second here (more on a Linux SSD). These figures were measured on the first version of this PR and were not re-measured since: the steady-state append path has not gained an fsync (the pending-fsync bookkeeping is an in-memory lookup, and the retry-after-failure directory fsync only runs after a failure).
Known limits
iterate,statsand prune).fsyncRfc64DirectoryV1has no soft-fail off Windows, so a filesystem that rejects a directory fsync (EINVAL/ENOTSUP) now makespersistMeta, and thereforeappend, reject. The RFC-64 durable stores already require the same of the node's data directory.renamethat rejects yet took effect (taken as atomic), and the unlinks ofinit()'s sweep of orphan logs, corrupt metas and stale temps (nothing acknowledges them; a resurrected orphan is reaped again by the next init)..metaor the log (EIO, EACCES, EMFILE) rejects the load and the mutation that needed it (nothing cached or persisted), andprune()keeps sweeping past a CG it cannot load and rethrows the first failure.devnet.sh's stop phase still signals anodeoroxigraph*process that mentions the node's home directory (for example aDKG_HOME=<home> node cli.js status), as it always has; only other processes that merely name the home are now left alone. The ownership check cannot tell apart a recycled PID that became another daemon of this same checkout, and PIDs are verified once and signalled later (only liveness is re-checked at the signal).>=to>in the retention plan.CHANGELOG.mdentry: the file has no[Unreleased]section at this base (the top section is the shipped 10.0.21).Test Plan
packages/agent/test/swm/host-mode-store-{durable-writes,tail-recovery,cold-init,dirsync-retries}.test.ts(new: the former single 64-test durability suite split into four files with a sharedtest/_helpers/host-mode-store-durability-support.ts, the same 73 test names including the nine unreadable-file tests added since; registered invitest.unit.config.ts): exact open / fsync / rename / directory-fsync sequence for append, meta and prune; rename, directory-fsync, write, file-fsync and frame-fsync failures; temp cleanup and handle close; cache drop on failure; best-effort reconcile; stale-temp sweep, decoys and the live-temp guard; torn-tail repair (partial payload, partial header, once per process, re-check after a failed append, uninspectable tail); durable cursor ahead of a lost frame; slow cold load vs a concurrent append; readers racing a prune; single-owner cold initialization (a delayed reconcile cannot restore a cleared flag, every mutator and append waits for a paused cold load, concurrent cold loads share one meta read / log scan / reconcile write, two loads in one tick share one initialization, a rejected initialization is retried, a non-string id is a swallowed rejection); retry after a failed directory fsync (table over all four mutators, single and double failure, prune rewrite and prune unlink, reconcile write, covering fsyncs, in-flight and overlapping fsyncs, a target changed again while a covering fsync is in flight, three targets, a cap-eviction unlink).packages/agent/test/swm/host-mode-store-dirsync-model.test.ts(new, 1 test of 16 seeds, about 0.1 to 0.4 s;HOST_STORE_MODEL_SEEDSwidens it): a seeded model that holds the mocked directory fsync in flight, settles calls out of order and fails some, and checks every acknowledgement (the four marks, append, prune including unlinks) against a monotonic durable viewcd packages/agent && npx vitest run --config vitest.unit.config.ts test/swm/host-mode-store.test.ts test/swm/host-mode-store-durable-writes.test.ts test/swm/host-mode-store-tail-recovery.test.ts test/swm/host-mode-store-cold-init.test.ts test/swm/host-mode-store-dirsync-retries.test.ts test/swm/host-mode-store-dirsync-model.test.ts test/swm/host-mode-store-crash.e2e.test.ts test/swm/host-mode-key-canonicalization.test.tshost-mode-store.ts(31aff22), 57 of the 64 durability tests (several by timeout, since the base store lacks the features they wait for), the model test and all 9 crash e2e tests fail; the 7 durability tests that pass are deliberate regression pins (clean log untouched, uninspectable tail, lost-frame gap, lagging cursor, a failed initialization not poisoning the queued caller, best-effort reconcile persistence, non-string id). Against the earlier pushed head (6b881be), 34 of 64 and the model test fail. Changed-line coverage at this head: agent 272/276 (98.6%, gate 90%, base a94711b; the uncovered lines aredkg-agent-swm-host.ts:8205,host-mode-store.ts:378,host-store-meta-loader.ts:110,host-store-startup-reconcile.ts:47) on a run of 12 host-store and host catch-up suites, and cli 26/26, computed with the changed-line half ofscripts/ci/check-coverage.mjs(inspectCoverage+changedLinesFromDiff+changedCoverage; the script's whole-package floors cannot be met by a partial run, so it was not run end to end). At the round-3 head 28 mutations of the extracted modules (including the four below that the review asked for) each fail tests; one more (the init sweep reaping the log of a.metait could not read) failed none untilhost-mode-store.test.tsgained a test. The counts that follow were measured on the earlier single-file version: sharing the cold-load initialization (12 tests fail without it), not caching a rejection (2), installing the cache only after the reconcile write (7), keeping the reconcile write best-effort (3), the log-tail half of the cursor (19), keeping the initialization out of the write lock (behind it, every mutator test deadlocks), the pending-fsync mark and its completion on the early returns (mark*, prune's nothing-to-drop and no-log branches: 23 tests without the mark ever being cleared, 2 without the completion on the no-log return), the per-change generation (the old per-pathSet: 2 deterministic tests and the model), and the directory fsync after the prune unlink (6 tests and the model) each fail tests. The earlier-round mutations (removing the tail re-check, the seqno reservation, the live-temp guard, the log fsync, the directory fsync or the temp cleanup) were measured in earlier rounds, against the tests that existed then. Over 2000 seeds the model finds no violation in the current store and hundreds with the per-pathSet(372 violations in 259 seeds) or with no fsync after the unlink (997 in 296 seeds); the default 16-seed run caught the per-pathSetin 12 of 12 runs the implementer measured and 19 of 20 the reviewer did, so it is probabilistic and the deterministic tests are the gate.packages/agent/test/swm/host-mode-store-crash.e2e.test.ts(new, 9 tests) andtest/_helpers/host-mode-store-crash-child.ts: atsxchild runs a real store and SIGKILLs itself inside prune / persistMeta / append (mid-write, before rename, after rename, torn frame) by wrappingfs.promises; the parent reopens like a restarted daemon and asserts no torn tail, every acknowledged frame served, the cursor never backwards or reused, temps swept, and strict-greater-than paging to the end; every retained envelope is compared byte for byte with the child's fixture after every crash window (store read, paged catch-up read, raw file), and a mid-write / before-rename kill is checked to leave a torn / complete tempcd packages/agent && npx vitest run --config vitest.config.ts test/swm test/finalized-authority-host-share-paths.test.ts(package full config, every agent SWM suite)devnet/swm-host-store-durability(new suite, registered insuites.json,pnpm-workspace.yaml, the rootpackage.jsonand the lockfile importer): a 6-node devnet, hosting core node4 and curator edge node5 (config edits backed up and restored), a fresh curated CG; 5 xkill -9of the core the moment a frame lands, restart, assert on the core'sswm-host/files and on host-mode re-engagement; the curator pages the core withhost-catchupto the end from several cursorspnpm test:devnet:swm-host-store-durabilityand, as regression,pnpm test:devnet:core-peers-features7dadd698e(6 nodes): preflight verified all 12 PIDs of nodes 1-6; swm-host-store-durability 42/42 (3 live + 39 no-devnet, 138 s; 5 kill cycles); core-peers-features 6/6 (63 s, core-fill passing). One change touches the devnet after that session:devnet.sh restart-node's process-table sweep is limited tonodeandoxigraph*processes (66d411143, and412f315b7for a checkout path with a space). Its tests run the real script (a realtail -f <home>/daemon.logsurvives the stop phase andrestartNodeAndWait), but it has not been re-run on a live devnet. The output tail below is from the earlier session, before the module split. Caveat: a process-level kill does not lose page-cache writes, so the suite is a regression and soak check on a real daemon, not a base-versus-fix discriminator; the deterministic base-failing evidence is in the unit and e2e rows. (An earlier version of the suite, with the weaker sorted/unique no-reuse assertion, also passed against the base store build, 3/3 with the basehost-mode-store.jsswapped intodist; that was not re-run against the current assertions.) The no-reuse check itself was shown to fail on a real daemon: with adistwhose recovery drops the last complete frame on the first append after a restart, the suite fails in cycle 1 ("frame #4 was seqno 4 at the kill but is seqno 5 after recovery"), a log the earlier assertion accepts by construction (run at 7d5e55045; the check has not changed since).pnpm test:devnet:manifest;pnpm run test:scripts;pnpm run lint;pnpm --dir packages/agent exec tsc --noEmit -p tsconfig.jsonpnpm test:inventoryverified 2265 test files; test:scripts 497/497 pass;node scripts/audit-file-size.mjspasses; lint clean; agenttscclean; a throwawaytscover the changed devnet files reports only the pre-existingdevnet/_bootstrap/harness.ts(605)HDNodeWalleterror, which this PR does not touchDevnet output
Four of the five kills land in the append-to-cursor window (the cursor is one behind the log tail at the kill) and one inside a durable meta write (it left a temp file, which the restart swept). After each restart every frame that was complete at the kill is still in the recovered log, byte for byte, the cursor is recovered from the log tail, and every frame appended after the restart has a seqno above
max(cursor, last complete frame)at the kill. Against the base store build the same kills land "between writes" (the base append is near-instant), which is why the suite is a regression check and not a discriminator.What a devnet core needs before it keeps any private SWM ciphertext in this release (the suite arranges all three and undoes them on exit): the RFC-64 kill switch on the core and the curator (in catalog mode the legacy SWM transport, and host-mode custody with it, is inert);
swmHostMode.stripCiphertext=falseon the core; and a curated graph that allowlists only the author (a private graph whose roster has a reachable peer besides the author is delivered point-to-point and never gossiped, OT-RFC-49 WS-A). Separately, after the core restarts, the edge-to-core link is relay-only and flaps, so live gossip does not reach the core again until the edge dials it directly; the suite does that after every restart. That last point is an observation about the devnet transport, unrelated to this change.test/swmrun; the full agent lane was not run)pnpm run build:packagesandpnpm --dir packages/cli run build:prepared)dkg start: covered by the devnet suite instead (see the table)Review follow-up (2026-10-02): a small addition to the catch-up endpoint
Four review findings on the live suite's catch-up scenario needed the endpoint to say more than a count and a cursor.
f4d77b1c6adds two optional body fields toPOST /api/shared-memory/host-catchup, the operator's endpoint for confirming what a hosting core holds:maxEntriesPerRound: ask each host for smaller pages. The agent and the wire request already supported it.includeEntries: per peer, the seqno and SHA-256 of every envelope that peer served, in order.Without them the request and the response are unchanged.
6b3f97b9cthen strengthens the scenario: with ingestion stopped the.metacursor must cover the whole log; paging uses a page a third of the log and must cross page boundaries; and the served envelopes are compared, in order, with the frames on disk (seqno and ciphertext digest). The two new checks are pure functions with no-devnet tests.At this head:
log-frames.test.ts25/25, the route tests 33/33, agent and CLI type-checks clean, and the live suite 3/3 on a six-node devnet (five kill cycles, each between the frame append and the cursor write; 20 frames served from 0 in four pages).Review follow-up (2026-10-06): current testnet-canary and the file-size ratchet
The branch merges testnet-canary at
a94711bf0. The new file-size audit failed CI onhost-mode-store.tsandmemory.ts, hence the module split above; the durability suite is split by concern; and the devnet helpers now check ownership on restarts as well as kills (a just-SIGKILLed daemon shows as(node)in state?Eon macOS, which an early version of the verified restart refused; fixed). The last two commits narrowdevnet.sh restart-node's own process sweep to node and managed store processes (the executable name decides; a path with a space falls back to whatps -o comm=reports). Not run as one:check-coverage.mjs agent --baseend to end (partial runs only, see the unit row).Related Issues
None filed.
🤖 Generated with Claude Code