Skip to content

release: merge v0.0.242 back to main - #464

Merged
justrach merged 41 commits into
mainfrom
release/0.0.242
Aug 6, 2026
Merged

release: merge v0.0.242 back to main#464
justrach merged 41 commits into
mainfrom
release/0.0.242

Conversation

@justrach

@justrach justrach commented Aug 6, 2026

Copy link
Copy Markdown
Owner

The 2026-08-06 batch, assembled and verified on release/0.0.242. Twelve PRs came in through integration PR #463; the changelog and --version what's-new commit sits on top.

Closes #416, closes #440, closes #445, closes #441, closes #411, closes #391, closes #444, closes #415, closes #418.

(Those keywords live here rather than on the individual PRs: those merged into release/0.0.242 rather than the default branch, so their own closing keywords never fired.)

Do not merge before the tag is cut, signed and notarized — that is the maintainer's step.

What shipped

Token economics, the batch's headline. MCP tool schemas load on demand (#416): a real 13-tool server's catalog drops 10,505 → 3,696 bytes, 64.8%, about 1,700 input tokens off every request, with servers under 4 KiB staying eager so small ones are byte-for-byte unchanged. Tool results above a threshold return a handle rather than their contents (#440), collapsing a three-cap stack whose middle layer made the outer one unreachable. The durable-transcript prompt line waits for the first compaction (#445), since before one the live context is a superset of the file.

Compaction durability. The session transcript is append-only (#441), correcting a premise shipped in #410: <name>.session.json is one JSON object rewritten in place, so compaction permanently discarded history. Both sides of the boundary now speak — notes-to-self before (#391, drawn from the existing landing-reserve ledger rather than a second one) and a durable-state note after (#411).

Engine separation. #429 batches 2 and 3: nine files lose their terminal imports, session_run.zig falls 597 → 450 lines, and the inventory found approvals.zig was on the ban list as a consent surface rather than for terminal coupling, so its pure half became harness_policy.zig.

Two things that were not what they looked like. The anyOf failure was never a flake (#444) — the build runner prints a failure report for successful steps with non-empty stderr, and cached test steps never re-run the tests, so "warm re-runs passed" was vacuous. And #441's Windows CI failure was one real defect, not five test bugs: File.stat on a write-only handle is ACCESS_DENIED there, so every Windows transcript was 0 bytes while macOS ran 1050/1050 green with the fix reverted.

Verification on the assembled branch

  • zig build test1116/1116, exit 0
  • scripts/eval-tier1.shgreen; ratchet bumped 994 → 1116 (it was already 49 behind before this batch)
  • E2E against a freshly built ReleaseSafe binary: test-pty-codex-ws.py, test-pty-overflow.py, test-spill-artifact.py, test-tty-release-on-exit.py
  • CI green on zig, windows and sdk

Four cross-branch breaks the integration caught

Each was invisible to every individual branch: a compile error from #411 removing an alias #445 still used; two leaf modules invented for one constant; a feature silently deleted by a clean merge (#440's GRAFF_TOOL_HANDLE_BYTES knob, when #429 moved the function containing it — no conflict, no error, no failing test); and a process-global that segfaulted three unrelated tests.

The last two now have structural guards rather than one-off fixes, each proven by break-and-revert: a knob that stops being parsed fails a named test, and the prompt funnel cannot arm a process-global from a test build (removing that gate reproduces the original segfaults).

Known follow-ups

#452 (no GC sweeper for .graff/tool-results/), #453 (a session rename leaves the transcript line naming a deleted file), #462 (the write-only-handle Windows defect is still live in playbook.zig and serve_events.zig — the latter overwrites the durable EventLog from offset 0, which #431's attach-replay acceptance depends on).

Not in this release: #429 batch 4 (agent.zig), which must be cut from this integrated result rather than from main, and #430/#431 behind it.

justrach and others added 30 commits August 6, 2026 18:10
… subagents (#419)

A side panel over the workspace board that answers "what are my agents
doing" without opening each conversation. buildAgentOverview derives one
orchestrator row per conversation plus a row per task tool_start, pairing
each with its tool_end to resolve running/completed/failed; followup
presence resolves needs-input. Rows sort attention-first
(needs-input > running > failed > idle), current conversation ahead of the
rest, newest first within a tier.

This is the client half of #419. The cheap-model session recaps land with
the engine-side recap events; the panel already renders the status chips
and the per-row detail line the recap will populate.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
The instrument that produced the +240 tok/call prompt-delta finding and the
three-cap-stack finding lived only in a session scratchpad, which made every
number it produced unreproducible the moment that session ended. It moves to
scripts/eval/live-ab/ largely intact, with the changes needed to survive a
clone:

- the two arms' binaries come from GRAFF_EVAL_BEFORE / GRAFF_EVAL_AFTER
  instead of hardcoded worktree paths, resolved lazily so running a single
  arm needs only that arm's variable, and failing with an actionable message
  rather than a traceback;
- model and timeout are overridable (GRAFF_EVAL_MODEL, GRAFF_EVAL_TIMEOUT),
  with one model per invocation so the two arms cannot silently diverge;
- one_run now WRITES .harness/settings.json into each run copy rather than
  relying on a committed dotfile. Disabling the companion MCP server is a
  measurement invariant (its schemas move the input-token number), and the
  repo .gitignore excludes .harness/, so a shipped copy would not have
  survived a clone at all;
- run residue (.git, __pycache__, .pytest_cache) is stripped from the
  fixtures, and runs/ + results.json are ignored as outputs.

RESULTS-2026-08-05-prime-batch.md is kept verbatim as the measured baseline
those findings came from.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
The two issues describe one mechanism from two directions (addressability
vs resumption) and are not separately implementable: resumption needs an
identity that outlives the spawn, and addressing needs somewhere for a
completed child's state to live. Built apart they would produce two id
schemes, two registries and two message paths.

Grounding the merge against the source turned up three things worth
recording:

- #392's premise is false. It assumes the worker transcript already
  persists under .graff/subagents/, making resume "plumbing, not new
  state". writeSubagentDetail persists a final report plus metadata as
  markdown, with no message history, no tool-call record and no provider
  state. Worker-history persistence is the first slice, not a detail.
- Identity partly exists already: background spawns carry a stable id and
  a mutex-guarded registry, but in memory only, so it dies with the
  process. And there are already TWO id schemes (cards.zig's sa-<ord>-<hash>
  and the background u32) that must be unified first or the registry keys
  on one and the artifacts on the other.
- The arc collides with #441, which excludes subagents from the
  append-only transcript by design while #392/#417 require exactly that
  history. Resolved as opt-in retention per spawn, so #441 keeps its
  default and gains one documented exception, and unretained workers keep
  today's behaviour byte for byte.

Steering unifies to one delivery call with a trigger flag rather than
Codex's two tools, keeping the queued-never-awaited invariant that stops
mutual sends between busy agents deadlocking.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
## What changed

`prompts.transcriptLineWanted(caps, compacted)` is now the single gate on the
#410 durable-transcript line, and `armSessionTranscript` takes the session's
compaction state alongside the capability answer. At startup that state is
false, so a fresh session composes exactly the base prompt and pays nothing.
`prompts.noteSessionCompacted` flips it once, at the compaction boundary, and
re-derives all four prompt variants from the agent's stored `sys_base` —
root-only, idempotent, best-effort, the same shape as
`playbook_glue.refreshRoot`. `agent_compact` calls it from the two places the
live history actually stops holding what it held: `compact()` once a summary
is installed, and `compactOrRecover()` once an emergency trim has dropped
messages.

The #421 doctrine paragraphs are untouched — they are behavioral rails and
stay unconditional. Only the transcript line moved.

## Why

- Problem/failure mode: the live before/after evals put the #421/#410 prompt
  additions at +960 chars, about +240 input tokens on EVERY api call at full
  capability. The transcript line's share of that bought nothing for the
  overwhelming majority of sessions, which never compact at all: before the
  first compaction the live context is a superset of the file the line names,
  so it can only point the model at wording it can already see, and compaction
  rewrites that file in place anyway.
- Reason for this approach: mutating a system prompt invalidates the cached KV
  prefix, so the injection had to ride an invalidation that already happens
  rather than introduce a second one. Compaction replaces the history sitting
  behind that prefix, which makes it the one free moment — and the exact moment
  the file starts holding something the window does not.
- Constraints or trade-offs: `compacted` is deliberately NOT a `Caps` field.
  `Caps` is what the tool catalog reports and `detectCaps()` is settled before
  the first request; this flips mid-session, and folding it in would falsify
  both of those and put a non-segment condition inside the predicate
  `composeBase`/`full()` short-circuits on. The two halves meet in one
  predicate instead. A resumed session also starts uncompacted on purpose: the
  file it loaded IS the context it has.
- Rejected alternatives: hooking `emergencyTrim()` itself rather than its
  `compactOrRecover` call site. Its other direct callers drive
  partially-initialized test agents, so reading `sub`/`sys_base`/`session_name`
  in there would be reading undefined memory in the suite.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
## What changed

`src/tool_handle.zig` is the one size contract for a tool result. At or below
the threshold (4 KiB, `GRAFF_TOOL_HANDLE_BYTES`) a result reaches the model
exactly as the tool produced it. Above it, the complete bytes are written under
`.graff/tool-results/` and the model gets a handle instead: a bounded preview,
that file's ABSOLUTE path, the byte count, and a one-line shape hint measured
from the payload (line count, or the top-level keys of JSON, only after the
document scans clean end to end). Meta tools never touch it — runTools applies
it to the external half of a batch.

This collapses three caps that used to sit in series:

- `agent_tools.persistToolResult` + `toolPreviewText` (2000 chars + a bare
  path) are REMOVED. They were the handle idea without the information that
  makes a handle usable: the model was told where the bytes were and nothing
  about what was in them, so it could not decide what to ask for without
  re-running the tool.
- `tools.bash_stdout_cap` is DEMOTED from a context cap to a capture ceiling
  and raised 128 KiB -> 1 MiB, and `codedb`'s 64 KiB result truncation is
  REMOVED outright (the const survives as the compact-read capture bound).
  Both existed to protect context by destroying bytes; the handle protects
  context harder, non-destructively. At 128 KiB the handle would have lied for
  exactly the case that motivated the issue — a 168 KB log, a quarter of it
  already gone before anything could be written down.
- `capOversizedToolOutputs` / `perOutputCap` / the #409 spill are KEPT, and
  narrowed to a backstop. `tool_handle.effectiveThreshold` clamps the tool-time
  threshold at or below `perOutputCap()`, so a result this process produced can
  no longer reach the send-time pass oversized. What still can is history this
  process did not produce — a session resumed from a pre-#440 build — which is
  precisely where deleting the pass would mean destroying bytes instead.

A capability-gated prompt segment (`prompt_text.tool_handle_note`, gate
`.local_tools`) tells the model the contract exists. It ships only when the
tools that produce a handle and the tools that can open one are both present:
under `--no-local-tools` a path is not something the session can act on.

## Why

- Problem/failure mode: the send-time spill (#409) was unreachable for the
  local toolset. Every `execTool` result was already reduced to 2000 chars at
  tool time, so nothing ever arrived at the ~136 KB send-time threshold — in 12
  live runs against a 168 KB log, including a forced `cat`, it never fired
  once. Meanwhile the bytes that mattered had already been destroyed twice
  over, by the 128 KiB bash cap and the 64 KiB codedb cap.
- Reason for this approach: a cap only rescues bytes at the moment of overflow.
  Handles-by-default means a large result never occupies context in the first
  place, and one result costs one threshold however large it is. The filesystem
  is the namespace and bash is the REPL, both of which graff already has, so
  this needs no kernel and no new tool.
- Constraints or trade-offs: the preview budget per oversized result roughly
  doubles (2000 -> 4096 bytes), which is the price of a head worth reading next
  to a marker worth acting on; it is constant, where the 128 KiB / 64 KiB /
  136 KiB numbers it replaces were not. A handle is written whole or not at
  all, bounded by a per-process budget, because a partial one would make the
  byte count a lie.
- Rejected alternatives: deleting the send-time cap as dead code. It is dead
  for anything runTools produced, and only for that; a resumed old session
  still carries outputs that never met this contract, and for those the
  alternative to the pass is an overflow.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
## What changed

`prompts.resetSessionCompacted` is the inverse of `noteSessionCompacted`, in
the same idiom: root-only, a no-op when the flag is already false, best-effort,
and it re-derives all four variants from the stored `sys_base`.
`commands_session` calls it from `/clear` (right before the `saveSession` that
empties the file) and from `/new` (after the `session_name` reassignment, so
the re-arm reads the conversation that exists now).

`commands_session.zig` sat at exactly the 600-line cap, so its two
`resetConversationSteering` tests moved verbatim to a `commands_session_test`
sibling — the same move `agent_eval_control_tests` made off `agent_tools.zig`
for the same reason — with the reachability hook in `test_hooks.zig`. The
helper is now `pub` for that. Net: the file drops to 578.

## Why

- Problem/failure mode: `g_session_compacted` is process-global and only ever
  went true, but one process hosts many root conversations. `/new` mints a
  fresh `session_name` and `/clear` empties the history and saves over the
  file, so in both cases the durable file again holds no more than the live
  window — precisely the state the line is not worth its tokens in. A user who
  compacted a long session and then hit `/new` paid the ~+240 tok/call for the
  rest of the process, pointing at a file with nothing to recover: the exact
  waste #445 exists to remove, reintroduced through a common flow.
- Reason for this approach: both doors already rewrite the conversation, so
  the reset rides an existing boundary rather than adding a new one, exactly
  as the forward direction rides compaction. The flag is a boundary, not a
  one-way latch: the new conversation earns the line back on its own first
  compaction, naming its own file.
- Constraints or trade-offs: `/clear` is the arguable one. It keeps the
  session name, so once #441's append-only `.transcript.jsonl` lands there
  would be an artifact surviving a `/clear` and worth naming. But the #410
  line names `.graff/sessions/<name>.session.json`, the resume artifact, and
  `/clear` genuinely empties that file one line later — so resetting is
  correct for the line as it exists today, and #441 can revisit it.
- Rejected alternatives: hooking `saveSession` on an empty history, which
  would have covered every present and future reset door without touching a
  file at the line cap, but couples saving to prompt state and fires from
  call sites that have nothing to do with a conversation ending.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
## What changed

The third door onto the same state. `commands_misc`'s `/resume` handler calls
`prompts.resetSessionCompacted` after it assigns the resumed `session_name`,
matching the ordering `/new` already uses. `commands_misc.zig` had room (566),
so no test extraction was needed this time.

## Why

- Problem/failure mode: `loadSession` replaces the live history with a COPY of
  the file it just read, so the #410 line is redundant for exactly the reason
  it is redundant on a fresh session — and worse than redundant if left armed,
  because it would still name the conversation the user resumed AWAY from. A
  process that compacted before the `/resume` kept paying ~+240 tok/call to
  point at a path belonging to a different conversation.
- Reason for this approach: same idiom and same call-site ordering as the
  other two doors, so all three read identically and a fourth is obvious to
  add. The resumed conversation earns the line back on its own first
  compaction, which is the whole contract of the flag.
- Constraints or trade-offs: the reset is deliberately unconditional on
  whether the resumed file is large. A resumed history near the context cap
  will compact almost immediately and re-arm; guessing ahead of that would
  reintroduce a second, independent prompt mutation, which is precisely what
  riding the compaction boundary avoids.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
Every connected MCP server used to put its whole `inputSchema` into the root
tool catalog at connect time, so a session paid for every schema of every
configured server on every request. Since #345 gave servers a user-level
`~/.codegraff/mcp.json` home that follows a user into every checkout, that is
a standing tax rather than a per-project choice: one companion server was
measured at +2,568 input tokens per call whether or not the model ever
touched it.

Two-phase exposure, the same progressive disclosure the `skill` tool already
uses for SKILL.md bodies:

  - Phase 1. A deferred tool is still REGISTERED - qualified name, its
    description capped to one line, and a placeholder schema - so the model
    knows the tool exists and can ask for it.
  - Phase 2. The new builtin `load_tool_schemas` returns the full JSON
    schemas for the tools (or the whole server) it is asked for and ENABLES
    them for the rest of the session. It is a meta tool, handled inline by
    the orchestrator, which both re-renders the catalog immediately and
    leaves the loaded-schema set with exactly one writer.
  - Calling a tool whose schema was never loaded is refused with the exact
    next action ("call load_tool_schemas with {tools: [...]} first"), never a
    bare unknown-tool error.

Consent is untouched, and deliberately so. Deferral is about context cost,
not permission. Nothing in the new module reads or writes approvals, and the
enforcement point sits INSIDE execToolInner - downstream of
agent_tool_gate.gateTool - so the consent prompt is still reached in exactly
the cases, and the order, it was reached in before. Deferral can only ever
subtract a call that consent already allowed; it can never add one, and it
never short-circuits or reorders an approval check.

The eager/lazy default is size-based, per server: a server whose tools cost
at most 4 KiB of description + serialized schema stays EAGER and behaves byte
for byte as it did before. 4 KiB is roughly 1,000 tokens, and below that the
load round trip re-emits about as much as deferral saved, so the indirection
only pays above it - which also means small servers keep working exactly as
today and only the expensive ones change. Pins either way:
`GRAFF_MCP_EAGER=<names>` (or `*` for the pre-#416 catalog), a per-server
`"eager": true` in .mcp.json / the global config, `GRAFF_MCP_SCHEMA_BUDGET`
to move the threshold, and `GRAFF_MCP_LAZY_SCHEMAS=0` to switch it off.
Loaded schemas are cached per session, so loading is once per session rather
than once per call.

Measured on the bundled 13-tool Smolify manifest, a real MCP server of the
same order as the one that provoked the issue: the catalog served for that
one server drops from 10,505 to 3,696 bytes - a 64.8% cut, roughly 1,700
input tokens off every request for the whole session. tool_schema_tests.zig
asserts the drop at a conservative 50%.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
#444 is not a flake and nothing was ever failing. `zig build test` exits 0
on every one of these runs; what it also does is print a step-failure tree
and a red `failed command: .../test --listen=-`, and the line people read as
the cause — `tool 'fake' has a JSON Schema 'anyOf'` — is a passing negative
control's expected diagnostic that merely sat in the same buffer.

The mechanism is in the build runner (zig 0.17.0-dev.813,
lib/compiler/Maker.zig:1164): a Run step whose `result_stderr` is non-empty
is handed to `printErrorMessages` REGARDLESS of its result, which renders
the failure tree, dumps the captured stderr, and — because the failed
command is recorded unconditionally before the child's term is known
(Maker/Step/Run.zig:2157) — prints `failed command:`. `printStepFailure`
even emits a bare `" w\n"` placeholder for this case, which is the
`+- run test w` line. Exit status stays 0 throughout.

So any stderr at all from a PASSING suite produces the whole alarming
block. This suite emitted twelve lines of it: the negative control's
`std.debug.print`, and eleven subagent activity lines — `[test]`,
`[workflow]`, `[diversity]` — which reach stderr because a test-built Agent
has no writer and `say()` falls through to the tick-gate path.

Why it looked correlated with the first run after a fresh compile: a zig_test
Run step sets `has_side_effects = false`, so `zig build test` with an
unchanged binary is a manifest cache hit that does not execute the tests at
all (`+- run test cached`, 0.15s). Only the first run after a compile ever
runs the suite, so only it can produce stderr. The "warm re-runs passed"
observation was vacuous — those runs ran nothing. And seed-independence
follows for free: there was no randomness involved, only presence or absence
of a cache hit. Reproduced 9/9 byte-identical (2250 bytes, exit 0) on
consecutive forced-recompile runs.

The fix makes the trigger structurally impossible rather than quieter:

- tick_gate gains `writeLine`, gated on `emit_to_stderr = !builtin.is_test`,
  a comptime constant. Every stderr write on the worker-line path now goes
  through it, so a test binary cannot reach one. The gate's accounting is
  untouched: offers are still held, drops still counted, released lines
  still counted.
- The four raw `std.debug.print("  [workflow]/[diversity] ...")` call sites
  now use the new `tick_gate.workerPrint`. They were always this module's
  business — its own doc names them — and bypassing it meant they neither
  honoured the line-boundary gate nor could be elided.
- tool_schema_tests mutes the offender diagnostic for the ONE call meant to
  fail. A real offender still prints, where stderr is exactly what you want.

Guarded by a new test asserting `!emit_to_stderr` at both comptime and
runtime while driving the real `workerLine` path, so removing the gate fails
the suite. `zig build test` now emits zero bytes; suite 1043 -> 1044,
tier 1 green.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
Corrects a false premise. #438 (issue #410) shipped a prompt line telling the
model that the session file preserves what compaction discards. It does not:
`.graff/sessions/<name>.session.json` is a SINGLE JSON object whose `messages`
array is rewritten in place, so the next autosave after a compaction drops the
pre-compaction history permanently. "Grep your own conversation log" therefore
had no graff equivalent — the trace and trajectory JSONLs record events for
telemetry, not the conversation for recall.

`.graff/sessions/<name>.transcript.jsonl` is that equivalent: one line per
message as it FIRST enters history, appended by a single positional write at
the end of file and never rewritten. A line is the message's provider-native
JSON verbatim (JSON escapes every control character, so a message is always
exactly one line), which makes the file greppable with no tooling and makes a
resume's re-seed exact — hashing a line reproduces the digest that wrote it.

Hooked at session.queueSave, the point where the autosave already observes a
new message, so none of the history's ~25 mutation sites needs a hook of its
own. Identity is a MULTISET of the digests on disk, not a position: compact()
builds `[handoff summary] ++ recent_messages`, putting the summary at the front
of the history while it is the last line in the file, and any position-based
match would re-append the retained tail behind it. Counting instead records the
summary alone. An ordinary turn takes a fast path (two serializations plus one
per new message) and only a rewrite, a resume, or a session's first append pays
for a full walk.

Bounds, as the issue asks:

  - Subagents excluded. Their history is never persisted, so there is no
    durable session to attach a transcript to — #409's rule, unchanged.
  - One lifecycle, not two. #409's sweep now reclaims transcripts and artifact
    dirs together (tool_spill.sweepSessionsOnce), by the same rule — the
    session file is the ground truth for "this session is gone" — and the same
    grace window. It now also runs at the first transcript append, so a run
    that never spills still collects what deleted sessions left behind.
  - The size cap ROTATES rather than head-truncates. Head-truncation means
    reading the file, dropping a prefix and writing what is left back over it:
    an in-place rewrite of the history, which is the exact failure mode this
    change exists to fix, and one a crash halfway through turns into total
    loss. A rename is one atomic syscall, the old bytes survive intact under a
    sibling name the model greps identically, and disk use is bounded at two
    generations (16 MiB each) rather than merely slowed.

Tests 1043 -> 1050. The rewrite case asserts the pre-compaction file is an
exact byte PREFIX of the post-compaction one and that the discarded detail is
still greppable; repeated autosaves over an unchanged history add nothing; a
resume re-seeds from disk instead of re-appending its restored history; a
subagent writes nothing; rotation leaves the previous generation byte-identical.
Verified end to end against a mock provider: two processes, four turns, four
lines, no duplicates.

session_transcript.activePath(root, arena) is the accessor #411's
post-compaction note should use — null for a subagent or a session with no
transcript, so the note can never cite a file that does not exist, with
lineCount() for its "N messages".

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
Batch 3 of #429: skills.zig, skill_docs.zig, review.zig, imagegen.zig and
agent_prompt.zig go from 1/2/1/1/2 banned imports to zero. Only one of the
seven was an actual terminal write, so the batch splits three ways.

Four of the seven were approvals.zig, and none of them wanted the approval
gate: skills and skill_docs want the settings-file path, imagegen wants
confinedPath/noSymlinkEscape, review wants readOnlyAllowed. approvals.zig is
already terminal-free — it is on the ratchet's list because it is the CONSENT
surface, the mutable allow-list a human grows by answering y/n. So the pure
half moves to harness_policy.zig (settings location, path confinement, the
command classifiers) and approvals.zig keeps the session state. A pure move:
every decl is re-exported unqualified, so method bodies and `Approvals.x` call
sites are untouched, and a new test asserts the re-exports ARE the aliases
rather than a second copy — a security predicate that drifts is a hole.

agent_prompt.zig is the real inversion. prompt() now gathers the session into
one `.prompt_ready` event and emits it; the palette, the badge frame and the
#209 width budget move to agent_prompt_render.zig behind TuiSink, the same
shape agent_stream_render/agent_tool_render have. One new variant and its
payload types:

  prompt_ready: PromptStatus  — what the status line SAYS: model, provider,
    cwd, the mode badges, and the meters as values (CostMeter distinguishes
    off / flat-rate / unpriced / a real figure; ContextMeter carries tokens,
    window and the compaction threshold). No widths, no colors, no assembled
    segments — which badges survive a narrow pane is a rendering decision.
    A presentation pulse: the wire has never had a prompt, because a --json
    client drives its own turns, and the emit site still returns early on
    json_mode rather than relying on a silent sink.

ReasoningEffort and PrivacyTier are the vocabulary's own enums rather than
main.zig's and learning_privacy's, so a transport-split sink can read an event
without importing the engine; the emit site maps with exhaustive switches and
a test pins the two sets in step.

skill_docs.zig gets a relocation, not an event pair, and deliberately: its
ansi use was the `/skills` command surface, whose other half (the companion
rows) is printed inline by commands_session.zig. Routing one half of one
command through a contract with no wire shape for it, leaving the other half
inline, would manufacture exactly the divergence #422 exists to remove. The
command layer is frontend territory, so handleRemove/handleAdd/printSection
move whole to skill_docs_render.zig and skill_docs.zig goes back to being the
skill subsystem its own header describes. commands_session.zig stays at 600
lines by repointing its existing import rather than adding one.

One behavioral note: promptLine swallows a write error where prompt() used to
propagate it, because a sink emit returns void. Every other TuiSink branch
already behaves that way, and the only path there is a terminal that vanished
mid-prompt.

Proven byte-identical by a golden before/after harness: a scripted-model PTY
session at three pane widths (160/60/34), capturing the status line with and
without usage, the `/skills` catalog, and remove/add/unknown — 40 artifacts,
raw bytes and rendered text, all identical between the pre- and post-change
binaries. Capture windows open only after a turn has settled, so no spinner
frame is inside one and every .raw is deterministic (verified by two runs of
the same binary agreeing). The layout is additionally pinned in-tree: the
renderer's tests assert the whole line byte-for-byte at four widths, including
the #209 shed order and the rule that the cache badge never floats free of the
context meter. tier 1 green; 1051 tests pass.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
Compaction hands the conversation to a summarizer, and a summarizer optimizes
for a readable account of what happened. The things a working agent cannot
cheaply re-derive - the exact line it was editing, the approach it already
ruled out, why it picked B over A - are precisely what a summary drops as
uninteresting. Codex solves this by reserving buffer tokens and firing one
prompt right before rollover; this is that, on graff's substrate.

When compaction is imminent the root now spends one bounded, tool-less turn
writing a note to its future self (subgoal, file:line anchors, decisions, dead
ends), stored outside the conversation and re-injected verbatim after the
history it describes is gone.

BUDGET: one ledger, not two. #390's landing reserve (phase_budget.Ledger) is
extended with cost_precompact_note and Ledger.affordsHarnessNote rather than
given a sibling. The note is the JUNIOR liability on that ledger - narration
is mandatory and owns the reserve, the note is optional and must fit ON TOP of
it, the same P3 gate judges pass. A second reserve would be a second ledger,
and two ledgers each holding back "the last call" double-count the same pool,
which is the bug #390 was filed for. The call then goes through the shared
RunBudget like any other, so it is counted where it is spent. A token buffer
(note_reserve_tokens) is the second half of the gate, checked against the
window. A test derives the call boundary from landingReserve() itself, so a
parallel reserve would move it and fail.

PERSISTENCE: a session-scoped append-only store (.graff/notes/<session>.notes.
jsonl) composed into the ROOT system prompt beside HARD CONSTRAINTS, not
playbook items with source=session. The playbook is project-scoped and
cross-session by design, caps items at 240 bytes, rides every subagent brief
through rideBrief, and keeps items until a user retires one by id - all four
are wrong for working state that is superseded by the next note and must never
reach a worker. What is borrowed is the MECHANISM, exactly: append-only JSONL
written whole-or-not-at-all, replayed deterministically, and assembled from the
file at injection time rather than from conversation memory. That last property
is the whole feature - there is no in-context copy for the next compaction to
paraphrase away, and the system prompt is re-sent verbatim on every request.

Never fires for a subagent (checked first, before anything is measured), at
most once per history generation however often compaction is retried (#379's
loop would otherwise buy a note per lap), and every refusal is a named skip:
compact() ignores the result, so a failed, empty or unwritable note costs
nothing but the note.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
The instrument that proved batch 3 behavior-preserving lived only in a session
scratchpad, which made its verdict unreproducible the moment that session
ended. Three separate issues depend on it — #429's remaining batches ride the
same bar, #430's acceptance is golden byte-identity in both modes, and #431's
attach-parity eval is this harness pointed at two transports — so it moves to
scripts/eval/golden/ in the shape scripts/eval/live-ab/ established:

- both arms' binaries come from GRAFF_EVAL_BEFORE / GRAFF_EVAL_AFTER, the same
  variables the live-ab harness uses, resolved lazily so capturing one arm
  needs only that arm's variable, and failing with an actionable message
  rather than a traceback. No path in the file is absolute: ROOT derives from
  __file__ and the shared PTY driver is imported through it;
- the determinism self-check is a first-class step and gates everything after
  it. The same binary is captured twice and must agree before any before/after
  diff is believed, because PTY capture has several ways not to be stable and
  a diff from an unstable instrument cannot be attributed;
- the session script is data (STEPS x WIDTHS), so the next batch adds a window
  rather than rewriting the driver;
- runs/ is gitignored. Only the instrument ships.

The README writes down the capture-window race that cost real time here: a
window opened at `cursor = len(raw)` straight after wait_for_literal can fall
on either side of readline's 12 setup bytes, and the failure is stable per run,
so one arm can look self-consistent while the two arms differ by a prefix with
byte-identical content behind it. Every step now settles before its window
opens. The known gaps are documented too rather than left to be rediscovered:
no NO_COLOR arm (this build draws no prompt under it), no --json arm (these
surfaces have no wire shape), and no approval prompt (#430's, not this).

Verified end-to-end on batch 3's two binaries: self-check green, 40 artifacts
byte-identical, exit 0. Also verified it FAILS correctly — a tampered artifact
and a deleted one are both reported, exit 1 — since an instrument that always
says "identical" is worse than none. zig build test green; tier 1 green.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
Batch 2 of #429: session_start, session_run, providers and hooks lose every
terminal-cluster import (3/5/1/2 -> 0/0/0/0). What they used to print is a
typed event now, rendered by session_render.zig — the lifecycle's terminal
half, sibling of agent_stream_render.zig and agent_tool_render.zig.

Eight new variants, each a moment the vocabulary had no way to say:

  session_notice          the shared shape, reused at eleven call sites so the
                          cluster does not need a variant per line. `tone` is
                          the only rendering choice delegated (it is about
                          meaning); `lead` carries a badge like "⚠ YOLO" that
                          is colored apart from the rest of its line.
  session_banner          the startup line. The key hints on it are frontend
                          knowledge and moved into the sink.
  worktree_entered        -w entered its scratch checkout; inline styling on
                          the path forces structure.
  saved_model_unavailable startup's twin of provider_fallback. Fields, not
                          prose: a frontend offering "use it anyway" needs them.
  mcp_consent_prompt      TRANSITIONAL — only the QUESTION is inverted; the
                          stdin read stays for #430.
  provider_fallback       the one durable variant, because this moment has
                          always been the wire's `model` event. The wire's
                          fixed note and the terminal's context note differ,
                          and that difference now lives at the one translation
                          point instead of at the call site.
  session_saved           the durable session file landed.
  run_finished            #396's terminal handoff: the engine says the run is
                          over, a terminal frontend hands raw mode back.

Two of those four files printed nothing at all — hooks.zig wanted only the
settings path and one Windows pipe peek, so those move to leaves
(harness_settings.zig, win_api.zig) that approvals.zig and term.zig re-export.
session_run's settings/theme phase does legitimately draw, so it moves whole
to session_settings.zig rather than pretending otherwise.

engine_sink.zig gains a writer-backed sink: this cluster runs before an Agent
exists and after it stops mattering, and a null writer (a one-shot, `acp`) is
a normal state, not an error. Its tests move to engine_sink_tests.zig — the
file was one line under the ceiling.

Line cap relieved where #422 said it would be: session_run 597 -> 450,
session_start 461 -> 465, providers 589 -> 586, engine_sink 528 -> 327,
term 564 -> 519. Nothing over 600.

Proven byte-identical: 17 scenarios run under a real PTY against a binary
built from main and one built from this branch, diffing RAW terminal bytes,
escape codes included — banner, worktree, YOLO badge, displays-on, both MCP
config reports, the consent prompt answered both ways, skipped servers,
companion auto-connect failure, approvals/hooks loaded, saved-model-unavailable
(blocked and not), both selftests, a fatal one-shot, --json, and a real turn
through the offline mock down to the session-saved line. All identical.
1052/1052 tests pass (1043 before, +9), tier 1 green, Windows cross-build clean.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
One genuine portability defect, not five test bugs. Every Windows failure —
four wrong line counts and one null deref — is the same root cause, and the
CI log names it: each count was "found 0", and activePath was null because
nothing had been recorded.

`appendWhole` opened the file write-only and then called `File.stat` on that
handle to find the append offset. On Windows, Io.Dir.createFile maps `read`
straight onto the NT access mask (`GENERIC = .{ .WRITE = true, .READ =
flags.read }`), so a write-only handle carries FILE_GENERIC_WRITE, which does
NOT include FILE_READ_ATTRIBUTES. `File.stat` is NtQueryInformationFile(.All),
which requires exactly that right, and std handles the resulting ACCESS_DENIED
explicitly. So on Windows the stat failed, `appendWhole` returned null having
ALREADY created the file, and every transcript was a 0-byte file: the feature
was inert on the platform, and #411's note would have cited an empty file.

`.read = true` on the create is the fix; a path-based `statFile` fallback keeps
a future std change from turning this back into a silent no-write rather than
an append. Nothing else about the write changed — it is still one positional
write at the end of file, still append-only.

This class is invisible to the POSIX suite: with the fix reverted, macOS still
runs 1050/1050 green. CI was the only possible signal, which is the argument
for the two hardenings below rather than for trusting a local pass.

The other two suspects, checked and reported rather than assumed:

  - Line endings: NOT a defect. '\n' is written explicitly and JSON escapes
    every control character, so a message is always exactly one line and no
    '\r' can appear. Hardened anyway, since the digests are over line bytes and
    a stray '\r' would silently re-append the entire history: `seed` trims a
    trailing '\r', and the tests now fail hard if a '\r' ever reaches the file.
  - Rotation's rename: NOT a defect. Windows rename does not overwrite, but
    Io.Dir.rename is the REPLACING variant on every OS (dirRenameWindows passes
    replace_if_exists=true; renamePreserve is the one that refuses a taken
    name). The rotation test only ever rotated onto a free name and could not
    have told the difference, so it now rotates twice and asserts the previous
    generation really was replaced.

INVARIANT, now stated where the paths are built (session_index.zig) because two
downstream branches depend on the answer: every `.graff/` path this harness
builds is forward-slashed on every platform, Windows included. Deliberately —
Windows accepts '/' in the paths reaching Io.Dir, and these strings are shown to
the model (#410's prompt line, #409's cap marker, #441's path inside #411's
note), so a shape that changes per platform buys nothing and costs goldens.
Separators were never the Windows failure here; the null deref was a downstream
symptom of the empty file. The corollary is for tests, and ee28d8c is the
precedent: assert on a basename or on a path built through the helpers, never by
matching a separator by hand. Both `activePath` assertions now go through the
accessor and check the cited file exists; the readers build their paths with
`transcriptPath`/`rotatedPath` so a test can no longer disagree with the code.

The tests moved to session_transcript_tests.zig: the hardening pushed the module
to 614 lines, over the ceiling. Reachability needed the `_ = ...` line in
test_hooks.zig's test block, not just the import — the count caught it at 1045
before it was added, and is 1050 again after.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
Compaction never said what the summary did NOT have to carry, so the model
treated it as total loss: it hoarded file contents and quoted tool output into
the summary, and the state the harness itself keeps came back only as its own
recollection of it, drifting a little more at every compaction.

Both halves of prime-agent's compaction design, in a new module
(compact_note.zig) so agent_compact.zig keeps its line count:

BEFORE. The summary REQUEST now carries a note saying what persists - every
file on disk, any goal/checklist (restated in full straight after the summary),
and #441's append-only transcript, cited by its real path and message count when
one is live. It asks for NAMES rather than contents and points the summary at
what disk cannot give back: decisions, dead ends, constraints, unfinished work.
The instruction itself still LEADS the request, so #379 classifies an empty or
truncated reply exactly as before.

AFTER. The new history head carries the harness's own ground truth, re-derived
at every compaction rather than copied forward - so a later summary that
paraphrases it away costs nothing, the next compaction regenerates it exactly.
Three fields, each omitted unless real: the files this session modified (from
/rewind's snapshot ledger, which is exact for write_file/edit_file/imagegen and
blind to bash - the note says so rather than passing a partial list off as the
whole diff); the #409 artifact paths the discarded messages' own spill markers
cite, so every handle named was written by a spill that succeeded; and the
transcript path via session_transcript.activePath, which returns null rather
than let the note invent a file. A subagent and a /review turn get nothing, and
a session with nothing durable to report keeps a byte-identical handoff.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
…name for #391

#391 independently created src/compact_note.zig for the OTHER side of the same
boundary: the buffer-reserved turn where the model writes notes to itself before
rollover. This module is the harness's own ground truth injected around the
handoff - which is what handoffMessage is and what the note actually rides - so
the name it now has says what the old one only implied, and a future reader
cannot confuse the two halves. #391 owns four files in that cluster, this owns
one, so the rename is the cheap side.

Pure rename plus the local alias (compact_note -> handoff_note) at its two call
sites and one test import. agent_compact.zig stays at 564, still net zero
against its branch point. Suite unchanged at 1056; the module keeps its
reachability through agent_compact.zig's production import, so no test_hooks
wiring changes.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
CI's codex PTY scenario caught the note turn. Investigating it turned up a
real gap and corrected a wrong assumption about where the boundary is.

THE GAP, now closed. #391 is about a PLANNED rollover: context is filling, so
reserve a buffer and spend one call before the window turns over. It is not
about a rescue. Once the provider has rejected a request for exceeding the
window (last_request_context_overflow), or the meter is at the destructive-
recovery boundary where compactOrRecover may drop real history
(Provider.nearContextLimit, 95%), the session has demonstrably run out of room
and a note is the last thing it can afford. decideRoom becomes decideContext
and refuses those outright, BEFORE the token-buffer question — which would
have said yes, because at 95% of a 200k window there are still 10k tokens of
nominal room. contextOf reads the same two signals compactOrRecover uses to
authorize destructive trimming, so the note fires exactly when compaction is
scheduled and never when it is damage control. Two unit tests, one pure and
one through maybeWrite with a live provider.

THE ASSUMPTION, corrected with evidence. run_midturn_compaction_scenario is
NOT the recovery path. test-pty-codex-ws.py sets the server meter to 90% of
the window and codex_ws_test.py says why: "Cross compact@ (80%) but stay below
the destructive recovery boundary (95%)." runTurn's mid-turn gate is
inputOverCompactThreshold (80%), and trim_on_fail is false there. So it is the
planned rollover, and the single most representative instance of #391 in the
tree — a long tool-loop turn crossing the threshold mid-flight, the failure
mode the issue was filed over. Suppressing the note there would suppress it on
the main path.

So the scenario is taught the new shape rather than the note suppressed, and
the #195 WS invariant is CHECKED rather than argued: the note turn is a quiet
tool-less SSE request that opens no WebSocket (connection_id is None), runs
inside runTurn's existing closeCodexWs bracket against a throwaway clone of
history, and leaves the choreography untouched — ws_connections stays 2, the
post-compaction turn still re-anchors on a fresh socket, and no request
carries previous_response_id. Verified by running it, not by reasoning.

Both compaction fixtures are re-keyed on the request's last USER TURN instead
of its ordinal. An ordinal-keyed fixture re-targets silently when compaction
gains a step: the summary reply landed on the note turn, the note reply became
the handoff summary, and midturn still went green while proving something
else. The transactional scenario failed outright for the same reason.

And the scenario now pins #391 end to end on the wire, which no unit test can
reach: the note is absent from `instructions` before it is written, present
VERBATIM in every request after it, and absent from `input` on the
post-compaction turn — state, not conversation, with nothing in the history
for the next compaction to paraphrase away. Confirmed load-bearing by
disabling the injection and watching it fail.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
…8-06-batch

# Conflicts:
#	src/test_hooks.zig
…-06-batch

# Conflicts:
#	src/agent_compact.zig
#	src/prompts.zig
…-08-06-batch

# Conflicts:
#	src/session_run.zig
…-08-06-batch

# Conflicts:
#	src/approvals.zig
#	src/engine_events.zig
#	src/engine_sink.zig
…ches could not see

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
…batch

The baseline was 994 and had been stale on main since well before this
batch (49 behind at the start of the day, 119 after). AGENTS.md makes it a
hand-bumped release-cut chore; this is that cut.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
justrach and others added 11 commits August 6, 2026 20:11
…posed

Both bugs were already fixed as instances. These are the structural guards
so the CLASS cannot recur, each proven by break-and-revert.

1. A knob that stops being parsed. #440 added GRAFF_TOOL_HANDLE_BYTES to
   setupSkillsAndTheme while #429 batch 2 moved that function to another
   file; git merged both cleanly and the knob simply ceased to exist. No
   conflict, no error, no failing test. The env parsing is extracted as
   session_settings.applyEnvKnobs, and session_settings_tests drives it with
   a recording stub that asserts every GRAFF_* name is actually QUERIED —
   not what it does, but that it is asked for at all, since a dropped block
   leaves nothing to assert against. Deleting that same knob's parse now
   fails with "knob 'GRAFF_TOOL_HANDLE_BYTES' is no longer read".

2. A process global that segfaults unrelated tests. #391's note store is a
   process global whose composition reads the filesystem through agent.io;
   #445's tests drive the funnel with stub Agents whose io is undefined, so
   once anything armed the store every later composition crashed inside
   compact_note.pathFor, three tests from the cause. #391's own tests were
   disciplined about this — the gap was the IMPLICIT arming inside
   setRootSystemPrompts. That is now gated on !builtin.is_test, the same
   shape #444 used for stderr: a test can only arm the store deliberately,
   and the funnel stays the pure string function the suite relies on. The
   store also now OWNS its name rather than borrowing a caller's buffer,
   with over-long names disarming rather than truncating, since a truncated
   name would read a DIFFERENT session's notes.

Removing the gate reproduces the original three segfaults, so the new
regression test pins the fix rather than merely describing it. The per-test
band-aids added while diagnosing were removed and the suite stayed green,
which is what proves the structural fix is carrying the weight.

1116/1116, tier 1 green, ratchet 1113 -> 1116.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
integration: the 2026-08-06 batch into release/0.0.242
The batch's entries: lazy MCP schemas (#416), handles by default (#440),
the append-only transcript and its Windows defect (#441), both halves of
the compaction boundary (#391, #411), the deferred transcript line (#445),
the anyOf non-flake (#444), engine output inversion batches 2 and 3 (#429),
the GUI agent overview (#419), and the two eval harnesses that moved out of
scratch.

cli.zig what's-new gains 0.0.242 and retires 0.0.229, the same shape the
0.0.241 cut used.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
#418 asked whether graff ever lets a subagent's tokens inflate its
parent's context-window estimate, which would fire compaction early on
fleet-heavy runs. Audited the four named surfaces and the answer is no,
by construction rather than by a save/restore dance:

  * the BILL is process-wide. Agent.recordCost writes pricing.g_cost, so
    a child's spend lands in the [usage] footer, /cost and the --json
    cost_usd without anyone forwarding it.
  * the METER is per-Agent. last_context_tokens / context_local_tokens
    are fields on the Agent that made the request; runSub builds the
    child its OWN Agent, and ToolCtx carries no pointer back to the
    parent, so a child has nothing to write to.
  * run_budget/phase_budget ledger model CALLS, not tokens. Neither has
    a token field to conflate.

A completed child hands back SubRun{output, usage}. The foreground path
(execSubagent) keeps output alone; the background path (agentJobPump ->
agent_output) stores usage on the job and renders it through
agentStatusText -- as text the model reads, never as a number the
parent's meter adds. The in-place layers that DO reuse the root Agent
(review.Context, compact()) already capture and pin the meter back.

Nothing to fix, so the test is the deliverable: two parents seeded
identically, one of them fleeted with eight workers that each end past
the parent's compaction threshold. Their meters stay equal field for
field, neither trips inputOverCompactThreshold, and the run tally moves
by exactly the children's tokens and USD. The second test pins the one
place a child's context total does cross over -- the report line -- to
the bytes it actually costs, not the 480k it names.

Suite: 1118 pass (was 1116). Tier 1 green.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
Found while auditing #418. behavior_upload.recordApi accumulates the
CALLER's meter into one turn-wide total, root and subagent alike, so a
turn that fanned out eight workers reports roughly nine windows' worth
under a field named context_tokens. Correct for upload, and never read
back by the harness -- but it is exactly the field a future reader
reaches for when they want "how full was the context on this turn", and
the name gives them no warning. Answering that from here would
re-introduce the bug #418 exists to prevent, one layer over.

Documented at the type (every field is a sum over the turn's agents), at
the field (the authoritative per-agent occupancy is
Agent.last_context_tokens / effectiveContextTokens, which is what the
compaction gates read), and at the accumulator (the summing is deliberate
and belongs on this side of the boundary).

Comments only. Suite still 1118 pass; tier 1 green.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
A throwaway question answered from the LIVE conversation with tools
disabled, rendered once, and then dropped. "what did that error actually
mean?" is a question ABOUT the work in progress, not a step of it, and
asking it the ordinary way costs the session two permanent messages plus
their re-send on every turn after.

The turn runs on a throwaway agent over a container-deep CLONE of the
root's history in an arena that dies with the call - compact_note_glue's
shape, for its reason: send-time normalization mutates `messages` in
place, so only a clone may be handed to it. `root.messages` is read once,
to clone, and written never.

That is what keeps it out of BOTH persistence paths, which are not the
same path: the session file is a rewrite of `root.messages`, while the
append-only transcript (#441) is an OBSERVATION of it taken inside
session.queueSave - and being append-only, a line written there could
never be taken back. Neither runs here, and neither could see the
exchange if it did: the side agent's messages live in an arena the root
holds no pointer to, and its `sub = true` makes `record` refuse it at the
first line.

It is still billed. The request goes through Agent.request like any
other, so recordUsage/recordCost bank it in the same `pricing.g_cost` the
[usage] footer and /cost read. Not keeping an exchange is not a reason to
hide what it cost.

Tools are off by construction: `request(null)` omits the `tools` field
entirely rather than sending an empty array, for all three wire formats,
and `text_only` keeps it tool-less if it is ever driven through runTurn.

Single-shot, deliberately - the issue names multi-turn side threads a
possible v2, and there is no side history here to continue.

Both halves ship: the `/btw` slash command and a `{"type":"btw"}` --json
control for the GUI, whose answer event carries `persisted:false` so a
client mirroring the stream can tell it apart from a real turn.
mainloop.zig was at the 600-line ceiling, so the JSON controls that never
become a turn moved to json_controls.zig, taking the per-request tool
knobs with them; mainloop ends up three lines shorter than it started.

Proof, not assertion. The unit tests drive the real `build` and compare
the session file and the transcript byte-for-byte across the exchange;
a deliberate leak was planted and the transcript comparison itself is
what caught it. A tier-2 case then runs the real binary against the
scripted model and checks the request AFTER the side question: it must
carry no trace of it, which is the only way it could reach either file.
That case also goes red when the leak is planted.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
test(budget): pin the usage-attribution invariant (#418)
CHANGELOG gains #415 and #418; cli.zig what's-new gains the /btw line.
Tier-1 ratchet 1113 -> 1125: the two branches each measured their delta
against the same 1116 baseline independently (+7 and +2), so the merged
total is their sum.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
…t settled

Neither existing harness crosses a session boundary: live-ab measures token
economics within a run, golden proves byte-identity of rendering. This one
builds durable state in one session and makes a FRESH process recover it,
which is the only way to test whether #441/#391/#411 actually work.

Four properties had to be there, three of them found the hard way:

- Tokens are erased from disk when printed, so the eval measures memory
  rather than re-reading. A fact still on disk is one any competent agent
  simply reads again.
- Steps are chained rather than numbered, because a predictable sequence
  gets batched: 14 predictable steps produced 9 api calls and 3 compactions,
  the same 14 chained produced 45 calls and 15 compactions.
- Compaction is forced with GRAFF_COMPACT_PCT. GRAFF_CONTEXT applies only to
  models whose window graff cannot look up, so against a known model it is
  silently ignored and the run compacts zero times while appearing to work.
- Scoring separates correct / anywhere / fabricated, which caught a run
  reporting 0/14 exact that had in fact retained all 14, and which keeps an
  agent that invents plausible values from scoring like one that recalls.

The OptMem comparison is recorded rather than vendored: the arm needs
LH_MEMO_BIN pointing at a binary the operator supplies.

Result: base recovered 12/12 cold with 0 fabricated; OptMem retained 8 of
the 8 it stored but exhausted its call budget at step 8 of 12, costing 2.3x
the calls per step. The base agent persisted by folding `| tee -a` into a
tool call it was already making, which is the #440 doctrine reached without
a subsystem.

Co-Authored-By: Codegraff <blackfloofie@codegraff.com>
chore(eval): cross-session memory eval + the OptMem evaluation
@justrach
justrach merged commit 0773a32 into main Aug 6, 2026
11 of 16 checks passed
@justrach
justrach deleted the release/0.0.242 branch August 6, 2026 16:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment