Repository navigation
Port schema-harness behavioral trace schema + predict–verify loop into codegraff trajectories #246
Description
Activity
- added 16 commits that reference this issue
on Jul 19, 2026 Status after the review-and-implementation pass on PR #252 (commits d9c8b86..fc7c744) and collector PR justrach/zigrepper#143:
Proposal item matrix
# Item Status Evidence 1 Task-agnostic 9-kind event schema Done All 9 kinds now emittable: 5 lifecycle kinds automatic; tool_started/tool_finished (with args), action_taken, text_delta under opt-in rich capture (GRAFF_BEHAVIOR_TRACE=full, local-only until the collector versions them). Local stream was previously dead in every real session (exclusive-create collision, fixed in d9c8b86) 2 Expectation field + mispredict Done (first producer live) recordExpectedAction/recordMisprediction wired into the eval-driven loop (--eval/--until): commitment before the scoring command, misprediction on nonzero exit or missed target (67fa0d8). Plan-drop semantics not yet modeled 3 Verifier as per-recipe hook Partial The --eval command is the pluggable verifier and now emits commitments/mispredicts; the stop-on-red-repair-resume contract is not implemented 4 Notes template per task class Not started 5 Recomputable scorer Done (invariant met) scripts/score_run.py recomputes audit + behavioral metrics + replay RHAE from events alone; CI round trip in scripts/test-behavior-trace.py. Not yet invoked by graff --eval itself 6 Feed scores into fleet/DGM Substrate ready, loop not closed Collector stores typed rows recomputable by one SQL query, harness_scores.run_id joins signed scores to behavioral runs, Wilson-LCB promote cron exists. No behavioral score is admitted into selection yet Validation criterion: met and exceeded. Converting the published ft09 trajectory into codegraff.behavior.v1 and re-scoring reproduces its 100.0 RHAE exactly; the non-saturating bp35 trajectory reproduces 93.51 exactly including per-level values, so the math is proven off the caps. Note for future consumers: real schema-harness traces carry a 10th kind (turn_fallback) beyond the documented nine.
Remaining for full closure: notes template (item 4), stop-on-red repair loop (item 3), behavioral score admission into fleet/DGM selection (item 6), per-segment text_delta, and collector-side versioned validation for the four rich kinds before they can upload.
Follow-ups tracked in #255 (shipped, f4d7c7c), #256 (shipped, 67fa0d8), #257 (shipped, fc7c744; windows-latest CI is the runtime arbiter for the .cmd launcher).
Completion update for release/v0.0.215 (90a365b) and backend justrach/zigrepper#146 (45cc02b):\n\n- All nine event kinds are accepted end to end. Metadata upload exposes only controlled tool classes, byte counts, timing/error flags, and call pairing; exact names, args, text, paths, prompts, and code require explicit content opt-in and land in the isolated 30-day table.\n- The eval path now binds an expectation to the verifier, stops immediately on red, drops the remaining plan, writes fixed CONFIRMED/GUESSES/HYPOTHESES/FACTS/HISTORY notes locally with 0600 permissions, re-injects them after compaction, and requires a fresh green verifier result before completion.\n- Tournament candidates are evaluated on the primary suite, selected correctness-first then tool economy, and only the winner sees a fresh hidden holdout. Manual promotion remains required.\n- Tool counts and successful-tool behavioral scores are recomputed from one clean closed rich trace, admitted to comparison/ranking, surfaced as aggregate evidence, and signed in grade receipt v3.\n- Local validation passed: Zig build/tests, learning E2E, tournament E2E, adapter tests, live behavioral audit/scorer/replay, and source line limits. Release CI is green.\n- Backend validation passed: TypeScript typecheck, 58/58 telemetry tests, 57/57 gateway tests, and the full Zig source/tool smoke workflow. Production migration 0013 and worker version 872bfc6c-8a8f-4cb2-91dd-46e38ec1208d are live. A synthetic metadata-only nine-event smoke trace stored all typed events with zero content rows.\n\nRemaining generalizations are deliberately narrower than this issue’s working eval/tournament path: per-segment intermediate text capture (currently root-turn bounded) and equivalent predict-verify adapters for non-eval task classes. The issue stays open for those follow-ups; the evolutionary harness path requested here is operational.
Implemented and released in v0.0.225.
All six items:
- Nine event kinds exactly as proposed (src/behavior_trace_types.zig:3-13), with a mandatory kind/seq/ts/run_id/schema envelope and a comptime collision guard (src/behavior_trace.zig:64-100).
- Expectation + mispredict:
recordExpectedAction/recordMisprediction(src/behavior_trace.zig:288, :309); a mispredict drops the prior plan and blocks completion (src/agent_eval_control.zig:10-22). - Verifier as a pluggable per-recipe hook — the harness runs the configured
evalcommand and forbids running it via bash (src/schema.zig:145). - Notes template with CONFIRMED / GUESSES / HYPOTHESES / FACTS / VERIFIER HISTORY, re-injected as append-only context so it survives compaction (src/eval_memory.zig:14-55).
- Recomputable stdlib scorer scripts/score_run.py, embedded in the binary at build.zig:101.
- Loop closed into fleet/DGM:
behavior_score_ppmdrives comparison and tournament selection (src/learn_comparison.zig:49-75, src/learn_tournament.zig:50).
One deliberate deviation worth recording: the rich kinds (tool args, model text) are opt-in behind
GRAFF_BEHAVIOR_TRACE=full(src/behavior_trace.zig:564) as a privacy default; the lifecycle stream is on by default. Documented in docs/behavioral-trajectories.md:282-292.
Summary
Port the behavioral trace schema and predict–verify control loop from the ARC-AGI-3 schema-harness trajectories into codegraff's
.graff/trajectories, so our runs record what the agent believed and why it changed its mind — not just how fast it was. Then close the loop by feeding recomputable scores into the fleet/DGM machinery so the harness can improve over time automatically.cc @yxlyx
Data source
Reference dataset: https://huggingface.co/datasets/schema-harness/arc-agi-3-schema-traces
50 ARC-AGI-3 gameplay trajectories (25 GPT-5.6 Sol @ 95.35% mean RHAE; 25 Claude Opus 4.8 / Fable 5 @ 98.98%, 183/183 levels). Each trajectory ships
run.json, a streamedevents.jsonl, sanitized session JSONL, per-level snapshots, the agent's ownnotes.md+world_model.py, and a stdlib-onlyscore_trajectories.pythat re-derives every score from events alone and verifies it against the manifests.A perfect run (
claude_fable_opus/claude-fable-5_max_ft09_100.0: 6 levels in 78 actions, 100% RHAE) was dissected for this analysis.What the winning traces do that we don't
1. Rich, recomputable event schema (9 kinds,
seq+ts)run_started,turn_started(full env state + legal actions),text_delta(streamed reasoning),tool_started/tool_finished(with args),turn_committed(plan + reason),action_taken(action + resulting state),model_mispredicted(predicted vs actual + surprise message),run_finished(outcome summary). The trace alone is sufficient to recompute the score — the scorer is the audit.2. Beliefs as files, re-injected every turn
The agent maintains
notes.mdfrom a fixed template — Action semantics (confirmed/guessed), Current level, Hypotheses to test, Confirmed facts — and the harness re-injects it into every user turn. Memory lives in the file, not the context window; hypotheses are explicitly labeled GUESS vs CONFIRMED so the agent never acts on speculation without saying so.3. Executable world model + backtest ratchet
The agent writes
world_model_vN.pyimplementingpredict(state, action) → next_state, snapshotted per milestone (snapshots/cleared_level_N.py). Arun_backtesttool replays all observed transitions through the model; the agent is not allowed to plan until the backtest is green over full history (observed progression: 3/3 → 10/10 → 24/24 → 43/43 → 44/44). Plans are simulated against the model (BFS) before any real action is committed.4. Mispredict = hard stop
The harness checks each executed action against the model's prediction. On mismatch it emits
model_mispredicted, drops the remainder of the committed plan, and forces model repair + re-backtest before acting again. ft09's entire exploration cost was 9 mispredicts ≈ 9 wasted actions out of 78. Every committed plan also carries a stated hypothesis (turn_committed.reason), making exploration deliberate.5. Honest scope assessment
state belief + expectation → act → verifier checks → surprise ? stop & repair : continue.Where codegraff stands today
.graff/traces/<run>.jsonl(src/trace.zig)startup_phase,api(ms, bytes, context_tokens, cache_read),tool(name, ms, result_bytes) — performance only.graff/trajectories/<run>.jsonlsession,recipe(provider/model/effort + prompt_sha/toolset_sha),prompt,turn(ms, task, context_tokens, model_calls, tool_errors, ok),turn_error— stats only, zero content.graff/eval-log.tsvmemotoolsrc/fleet.zig,examples/dgm_loop.pyProposal
state/action/expectas opaque JSON; onlykind/seq/tsmandatory. Extendsrc/trace.zig/ trajectory writer; keep the existingrecipeline.expect:; a contradicting observation logsmispredicted. Zero cost when unused; no simulation required.run_backtest, coding getszig build test/ the eval command; same stop-on-red-repair-resume contract, pluggable command.scripts/score_run.py-style stdlib verifier, mirroringscore_trajectories.py).Suggested phases
graff --eval).Validation
Replay the dissected ft09 trajectory through the new schema and prove the emitted events reproduce its published RHAE score (100.0) — round-trip auditability from day one.
References
claude_fable_opus/claude-fable-5_max_ft09_100.0(6 levels, 78 actions, 9 mispredicts, 100.0 RHAE)min(115, 100*(h/a)^2), weighted mean by level, completion cap → RHAE