Skip to content

Port schema-harness behavioral trace schema + predict–verify loop into codegraff trajectories #246

Description

@justrach

Summary

Port the behavioral trace schema and predict–verify control loop from the ARC-AGI-3 schema-harness trajectories into codegraff's .graff/trajectories, so our runs record what the agent believed and why it changed its mind — not just how fast it was. Then close the loop by feeding recomputable scores into the fleet/DGM machinery so the harness can improve over time automatically.

cc @yxlyx

Data source

Reference dataset: https://huggingface.co/datasets/schema-harness/arc-agi-3-schema-traces

50 ARC-AGI-3 gameplay trajectories (25 GPT-5.6 Sol @ 95.35% mean RHAE; 25 Claude Opus 4.8 / Fable 5 @ 98.98%, 183/183 levels). Each trajectory ships run.json, a streamed events.jsonl, sanitized session JSONL, per-level snapshots, the agent's own notes.md + world_model.py, and a stdlib-only score_trajectories.py that re-derives every score from events alone and verifies it against the manifests.

A perfect run (claude_fable_opus/claude-fable-5_max_ft09_100.0: 6 levels in 78 actions, 100% RHAE) was dissected for this analysis.

What the winning traces do that we don't

1. Rich, recomputable event schema (9 kinds, seq+ts)

run_started, turn_started (full env state + legal actions), text_delta (streamed reasoning), tool_started/tool_finished (with args), turn_committed (plan + reason), action_taken (action + resulting state), model_mispredicted (predicted vs actual + surprise message), run_finished (outcome summary). The trace alone is sufficient to recompute the score — the scorer is the audit.

2. Beliefs as files, re-injected every turn

The agent maintains notes.md from a fixed template — Action semantics (confirmed/guessed), Current level, Hypotheses to test, Confirmed facts — and the harness re-injects it into every user turn. Memory lives in the file, not the context window; hypotheses are explicitly labeled GUESS vs CONFIRMED so the agent never acts on speculation without saying so.

3. Executable world model + backtest ratchet

The agent writes world_model_vN.py implementing predict(state, action) → next_state, snapshotted per milestone (snapshots/cleared_level_N.py). A run_backtest tool replays all observed transitions through the model; the agent is not allowed to plan until the backtest is green over full history (observed progression: 3/3 → 10/10 → 24/24 → 43/43 → 44/44). Plans are simulated against the model (BFS) before any real action is committed.

4. Mispredict = hard stop

The harness checks each executed action against the model's prediction. On mismatch it emits model_mispredicted, drops the remainder of the committed plan, and forces model repair + re-backtest before acting again. ft09's entire exploration cost was 9 mispredicts ≈ 9 wasted actions out of 78. Every committed plan also carries a stated hypothesis (turn_committed.reason), making exploration deliberate.

5. Honest scope assessment

  • What improves automatically: the agent's artifacts, within one run (the ratchet).
  • What improves manually: the scaffold (prompt/tools/scorer), across runs, by humans reading traces. The dataset + re-scorer make that loop measurable.
  • What never changes: model weights.
  • Portability: the event schema, notes pattern, and predict–verify loop are task-general; the world-model/BFS simulation is ARC-specific. For coding tasks, the compiler/tests/linter are the verifier — no simulation needed. General form: state belief + expectation → act → verifier checks → surprise ? stop & repair : continue.

Where codegraff stands today

Piece Location Records
Perf trace .graff/traces/<run>.jsonl (src/trace.zig) startup_phase, api (ms, bytes, context_tokens, cache_read), tool (name, ms, result_bytes) — performance only
Trajectory .graff/trajectories/<run>.jsonl session, recipe (provider/model/effort + prompt_sha/toolset_sha), prompt, turn (ms, task, context_tokens, model_calls, tool_errors, ok), turn_error — stats only, zero content
Eval log .graff/eval-log.tsv one TSV line per eval; write-once, not recomputable
Scratch memory codedb-pro memo tool findings/plans/errors; unstructured, not re-injected per turn
Fleet / DGM src/fleet.zig, examples/dgm_loop.py prompt-variant tournaments + eval selection — the self-improvement machinery schema-harness lacks
Capability schema-harness codegraff Gap
Tool calls with args ✅ ❌ (name+ms only) 🔴
State at each turn ✅ full ❌ 🔴
Plan + reason at commit ✅ ❌ 🔴 biggest miss
Expectation/mispredict events ✅ ❌ 🔴
Notes re-injected per turn ✅ templated 🟡 memo, unstructured 🟡
Run outcome event ✅ ❌ 🔴
Score recomputable from events ✅ stdlib scorer ❌ 🔴
Perf telemetry ❌ ✅ excellent 🟢 we win
Recipe reproducibility (shas) 🟡 ✅ 🟢 we win
Auto scaffold improvement ❌ manual ✅ fleet/DGM exists 🟢 we win

Proposal

  1. Task-agnostic event schema — adopt the 9 event kinds in the trajectory writer; state/action/expect as opaque JSON; only kind/seq/ts mandatory. Extend src/trace.zig / trajectory writer; keep the existing recipe line.
  2. Expectation field, not world models — any committed action may carry expect:; a contradicting observation logs mispredicted. Zero cost when unused; no simulation required.
  3. Verifier as per-recipe hook — games get run_backtest, coding gets zig build test / the eval command; same stop-on-red-repair-resume contract, pluggable command.
  4. Notes template per task class — CONFIRMED/GUESS/HYPOTHESES/FACTS sections, recipe-configurable, memo-backed so it survives compaction; strict mode for eval/benchmark runs, opt-in for interactive.
  5. Recomputable scorer interface — per-task metric, one invariant: score must be recomputable from events alone (a scripts/score_run.py-style stdlib verifier, mirroring score_trajectories.py).
  6. Close the loop — feed recomputed behavioral scores into fleet/DGM selection so scaffold variants evolve on decision-quality traces, not just perf. This is the "harness improves over time" piece schema-harness never built.

Suggested phases

  • Phase 1: items 1–2 (schema + expectation/mispredict) — the foundation everything hangs on.
  • Phase 2: items 3–4 (verifier hook + notes template).
  • Phase 3: item 5 (re-scorer wired into graff --eval).
  • Phase 4: item 6 (fleet/DGM integration).

Validation

Replay the dissected ft09 trajectory through the new schema and prove the emitted events reproduce its published RHAE score (100.0) — round-trip auditability from day one.

References

Activity

  1. justrach commented on Jul 21, 2026

    @justrach
    OwnerAuthor

    Status after the review-and-implementation pass on PR #252 (commits d9c8b86..fc7c744) and collector PR justrach/zigrepper#143:

    Proposal item matrix

    # Item Status Evidence
    1 Task-agnostic 9-kind event schema Done All 9 kinds now emittable: 5 lifecycle kinds automatic; tool_started/tool_finished (with args), action_taken, text_delta under opt-in rich capture (GRAFF_BEHAVIOR_TRACE=full, local-only until the collector versions them). Local stream was previously dead in every real session (exclusive-create collision, fixed in d9c8b86)
    2 Expectation field + mispredict Done (first producer live) recordExpectedAction/recordMisprediction wired into the eval-driven loop (--eval/--until): commitment before the scoring command, misprediction on nonzero exit or missed target (67fa0d8). Plan-drop semantics not yet modeled
    3 Verifier as per-recipe hook Partial The --eval command is the pluggable verifier and now emits commitments/mispredicts; the stop-on-red-repair-resume contract is not implemented
    4 Notes template per task class Not started
    5 Recomputable scorer Done (invariant met) scripts/score_run.py recomputes audit + behavioral metrics + replay RHAE from events alone; CI round trip in scripts/test-behavior-trace.py. Not yet invoked by graff --eval itself
    6 Feed scores into fleet/DGM Substrate ready, loop not closed Collector stores typed rows recomputable by one SQL query, harness_scores.run_id joins signed scores to behavioral runs, Wilson-LCB promote cron exists. No behavioral score is admitted into selection yet

    Validation criterion: met and exceeded. Converting the published ft09 trajectory into codegraff.behavior.v1 and re-scoring reproduces its 100.0 RHAE exactly; the non-saturating bp35 trajectory reproduces 93.51 exactly including per-level values, so the math is proven off the caps. Note for future consumers: real schema-harness traces carry a 10th kind (turn_fallback) beyond the documented nine.

    Remaining for full closure: notes template (item 4), stop-on-red repair loop (item 3), behavioral score admission into fleet/DGM selection (item 6), per-segment text_delta, and collector-side versioned validation for the four rich kinds before they can upload.

    Follow-ups tracked in #255 (shipped, f4d7c7c), #256 (shipped, 67fa0d8), #257 (shipped, fc7c744; windows-latest CI is the runtime arbiter for the .cmd launcher).

  2. justrach commented on Jul 23, 2026

    @justrach
    OwnerAuthor

    Completion update for release/v0.0.215 (90a365b) and backend justrach/zigrepper#146 (45cc02b):\n\n- All nine event kinds are accepted end to end. Metadata upload exposes only controlled tool classes, byte counts, timing/error flags, and call pairing; exact names, args, text, paths, prompts, and code require explicit content opt-in and land in the isolated 30-day table.\n- The eval path now binds an expectation to the verifier, stops immediately on red, drops the remaining plan, writes fixed CONFIRMED/GUESSES/HYPOTHESES/FACTS/HISTORY notes locally with 0600 permissions, re-injects them after compaction, and requires a fresh green verifier result before completion.\n- Tournament candidates are evaluated on the primary suite, selected correctness-first then tool economy, and only the winner sees a fresh hidden holdout. Manual promotion remains required.\n- Tool counts and successful-tool behavioral scores are recomputed from one clean closed rich trace, admitted to comparison/ranking, surfaced as aggregate evidence, and signed in grade receipt v3.\n- Local validation passed: Zig build/tests, learning E2E, tournament E2E, adapter tests, live behavioral audit/scorer/replay, and source line limits. Release CI is green.\n- Backend validation passed: TypeScript typecheck, 58/58 telemetry tests, 57/57 gateway tests, and the full Zig source/tool smoke workflow. Production migration 0013 and worker version 872bfc6c-8a8f-4cb2-91dd-46e38ec1208d are live. A synthetic metadata-only nine-event smoke trace stored all typed events with zero content rows.\n\nRemaining generalizations are deliberately narrower than this issue’s working eval/tournament path: per-segment intermediate text capture (currently root-turn bounded) and equivalent predict-verify adapters for non-eval task classes. The issue stays open for those follow-ups; the evolutionary harness path requested here is operational.

  3. justrach commented on Jul 31, 2026

    @justrach
    OwnerAuthor

    Implemented and released in v0.0.225.

    All six items:

    1. Nine event kinds exactly as proposed (src/behavior_trace_types.zig:3-13), with a mandatory kind/seq/ts/run_id/schema envelope and a comptime collision guard (src/behavior_trace.zig:64-100).
    2. Expectation + mispredict: recordExpectedAction/recordMisprediction (src/behavior_trace.zig:288, :309); a mispredict drops the prior plan and blocks completion (src/agent_eval_control.zig:10-22).
    3. Verifier as a pluggable per-recipe hook — the harness runs the configured eval command and forbids running it via bash (src/schema.zig:145).
    4. Notes template with CONFIRMED / GUESSES / HYPOTHESES / FACTS / VERIFIER HISTORY, re-injected as append-only context so it survives compaction (src/eval_memory.zig:14-55).
    5. Recomputable stdlib scorer scripts/score_run.py, embedded in the binary at build.zig:101.
    6. Loop closed into fleet/DGM: behavior_score_ppm drives comparison and tournament selection (src/learn_comparison.zig:49-75, src/learn_tournament.zig:50).

    One deliberate deviation worth recording: the rich kinds (tool args, model text) are opt-in behind GRAFF_BEHAVIOR_TRACE=full (src/behavior_trace.zig:564) as a privacy default; the lifecycle stream is on by default. Documented in docs/behavioral-trajectories.md:282-292.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions