Repository navigation
Define trace/trajectory file ownership and preserve JSONL integrity across concurrent processes #185
Description
Activity
Investigated end-to-end (against current
main). The three mechanisms in the report are all confirmed present. Here's what's implementable now vs. blocked.Trace — solvable (unique-per-writer file, model 2)
Give the trace a per-process name
harness.trace.<12-hex>.jsonl(freshio.randomid per process, same source as the scorerun_id— no wall clock). This eliminates the cross-process truncation and gives provenance in the filename. It's fully wireable: the one place that surfaces the fixed name to the agent (theharness.trace.jsonlmarker inmain_system_prompt,src/prompts.zig) is rewritten at prompt-build time to the real path, and/trace+ the startup banner print the actual file. Reader-facing change: consumers globharness.trace.*.jsonlinstead of a fixed name (same family, identical format). A prototype patch exists and passeszig build test(176).Trajectory — BLOCKED on the Zig 0.16 std
The shared-offset race (
writer.pos = st.sizethen a positioned write) needs an atomic append. Zig 0.16'sIo.Fileexposes noO_APPENDequivalent —Io.Dir.CreateFileOptions/OpenFileOptionsoffer onlyread/truncate/exclusive/advisorylock/permissions/resolve_beneath, and the POSIX backend never setsAPPEND. The only primitive is a whole-handle advisory flock, which for a long-lived accumulating file would serialize entire concurrent sessions, not individual appends. So atomic cross-process append is not achievable without either a std change (addO_APPEND) or a bespoke locked-append writer that fights the bufferedFile.Writer's ownpostracking. Not doing that fragile thing here.Net
- Acceptance criteria met by the trace fix: single documented policy, no cross-process truncation, complete-record + zero-bytes-on-failure preserved, provenance that survives restarts (not PID-based).
- Still open/blocked: trajectory atomicity (std limitation), startup detect/repair/quarantine of a torn final record, and the restart / forced-termination / multi-process-stress tests (the trajectory line is the blocker for the stress criterion).
Keeping this open. The trace half can land independently if we accept the glob-instead-of-fixed-name reader change; the trajectory half wants an upstream
O_APPEND(or a documented decision to accept per-epoch trajectory files instead of one accumulating file).Resolved on main by the run-scoped file split in
5aed7f1(#246).src/trace.zig:1-8:Every invocation writes its own
.graff/traces/<run-id>.jsonl(performance telemetry) and.graff/trajectories/<run-id>.jsonl(DGM archive nodes plus the experimental behavioral lifecycle/belief stream), so independent graff processes never truncate or seek/write through the same file. Every record is stamped with the run id, process id, and runtime session id before the event's own fields.That answers both halves of this issue directly:
- Ownership — each file is owned by exactly one process for its lifetime, keyed on run id (
tracePath/trajectoryPathatsrc/trace.zig:43-49). There is no shared cursor to race on, so the startup-truncation and same-append-position failure modes are structurally gone rather than merely locked around. - Integrity — records stay pre-serialized and flushed whole, and now additionally carry run id + pid + session id, so a reader can attribute every line.
Readers were updated to match: the archive reader scans all
.graff/trajectories/*.jsonlfiles rather than assuming one (src/trace.zig:336).Reopen if a shared-file path still exists somewhere I have not spotted.
- Ownership — each file is owned by exactly one process for its lifetime, keyed on run id (
Motivation
The existing JSONL fix prevents partial serialization through one in-process writer, but it does not protect shared trace and trajectory files from independent Codegraff processes.
Trace startup can truncate another process's output, while trajectory writers can observe the same file size and choose the same append position. Process-local mutexes do not coordinate these operations.
Current behavior / evidence
Complete records are pre-serialized, written, and flushed:
src/trace.zig:18-34Trace and trajectory synchronization is process-local:
src/trace.zig:41-47src/trace.zig:86-94src/trace.zig:117-175The trace is opened with truncating behavior, while trajectory append initializes its position from a separate file-size observation:
src/session_start.zig:180-201Existing visible coverage uses one in-memory writer rather than files or real processes:
src/trace.zig:221-235Proposed behavior
Select and document one supported ownership model for both trace and trajectory output:
Silent truncation, shared-offset races, and byte interleaving are unacceptable.
After choosing the model, add only the provenance needed to partition and correlate records unambiguously. This may be a writer/epoch identifier in records, headers, or filenames; do not require redundant per-record identifiers if ownership already supplies equivalent provenance.
Define behavioral crash guarantees. Do not equate a single
writecall with universal crash-safe atomicity; permitting and repairing or quarantining one torn final record is acceptable if documented.Acceptance criteria
harness.trace.jsonlandharness.trajectory.jsonl.Related issues