Skip to content

Define trace/trajectory file ownership and preserve JSONL integrity across concurrent processes #185

Description

@yxlyx

Motivation

The existing JSONL fix prevents partial serialization through one in-process writer, but it does not protect shared trace and trajectory files from independent Codegraff processes.

Trace startup can truncate another process's output, while trajectory writers can observe the same file size and choose the same append position. Process-local mutexes do not coordinate these operations.

Current behavior / evidence

Complete records are pre-serialized, written, and flushed:

Trace and trajectory synchronization is process-local:

The trace is opened with truncating behavior, while trajectory append initializes its position from a separate file-size observation:

Existing visible coverage uses one in-memory writer rather than files or real processes:

Proposed behavior

Select and document one supported ownership model for both trace and trajectory output:

  1. coordinated multi-process append;
  2. unique per-writer/per-epoch files; or
  3. exclusive ownership that fails safely when the destination is already owned.

Silent truncation, shared-offset races, and byte interleaving are unacceptable.

After choosing the model, add only the provenance needed to partition and correlate records unambiguously. This may be a writer/epoch identifier in records, headers, or filenames; do not require redundant per-record identifiers if ownership already supplies equivalent provenance.

Define behavioral crash guarantees. Do not equate a single write call with universal crash-safe atomicity; permitting and repairing or quarantining one torn final record is acceptable if documented.

Acceptance criteria

  • Trace and trajectory output use one documented ownership policy.
  • One live process cannot silently truncate another process's committed trace.
  • Independent writers cannot select the same trajectory offset or interleave record bytes.
  • Every committed record is one complete JSON object followed by one newline.
  • Supported concurrent writers can be partitioned and correlated unambiguously.
  • Provenance changes across writer lifecycles/restarts and does not depend on PID alone when IDs are required.
  • Existing readers remain compatible, or schema/version migration is documented and tested.
  • Startup detects and repairs or quarantines a permitted incomplete final record before appending.
  • A restart test verifies the retention policy and valid JSONL.
  • A forced-termination test verifies the documented crash boundary.
  • A real multi-process stress test verifies that every expected committed record appears exactly once and every retained line parses.
  • Tests cover both harness.trace.jsonl and harness.trajectory.jsonl.
  • Failure before serialization completes emits zero bytes, preserving the TUI: HttpConnectionClosing retries give up turn and corrupt trajectory log #86/PR Release 0.0.166: reliable JSONL logging (#86) + compaction-failure recovery (#88) #90 guarantee.

Related issues

Activity

  1. justrach commented on Jul 12, 2026

    @justrach
    Owner

    Investigated end-to-end (against current main). The three mechanisms in the report are all confirmed present. Here's what's implementable now vs. blocked.

    Trace — solvable (unique-per-writer file, model 2)

    Give the trace a per-process name harness.trace.<12-hex>.jsonl (fresh io.random id per process, same source as the score run_id — no wall clock). This eliminates the cross-process truncation and gives provenance in the filename. It's fully wireable: the one place that surfaces the fixed name to the agent (the harness.trace.jsonl marker in main_system_prompt, src/prompts.zig) is rewritten at prompt-build time to the real path, and /trace + the startup banner print the actual file. Reader-facing change: consumers glob harness.trace.*.jsonl instead of a fixed name (same family, identical format). A prototype patch exists and passes zig build test (176).

    Trajectory — BLOCKED on the Zig 0.16 std

    The shared-offset race (writer.pos = st.size then a positioned write) needs an atomic append. Zig 0.16's Io.File exposes no O_APPEND equivalent — Io.Dir.CreateFileOptions/OpenFileOptions offer only read/truncate/exclusive/advisory lock/permissions/resolve_beneath, and the POSIX backend never sets APPEND. The only primitive is a whole-handle advisory flock, which for a long-lived accumulating file would serialize entire concurrent sessions, not individual appends. So atomic cross-process append is not achievable without either a std change (add O_APPEND) or a bespoke locked-append writer that fights the buffered File.Writer's own pos tracking. Not doing that fragile thing here.

    Net

    • Acceptance criteria met by the trace fix: single documented policy, no cross-process truncation, complete-record + zero-bytes-on-failure preserved, provenance that survives restarts (not PID-based).
    • Still open/blocked: trajectory atomicity (std limitation), startup detect/repair/quarantine of a torn final record, and the restart / forced-termination / multi-process-stress tests (the trajectory line is the blocker for the stress criterion).

    Keeping this open. The trace half can land independently if we accept the glob-instead-of-fixed-name reader change; the trajectory half wants an upstream O_APPEND (or a documented decision to accept per-epoch trajectory files instead of one accumulating file).

  2. justrach commented on Jul 24, 2026

    @justrach
    Owner

    Resolved on main by the run-scoped file split in 5aed7f1 (#246). src/trace.zig:1-8:

    Every invocation writes its own .graff/traces/<run-id>.jsonl (performance telemetry) and .graff/trajectories/<run-id>.jsonl (DGM archive nodes plus the experimental behavioral lifecycle/belief stream), so independent graff processes never truncate or seek/write through the same file. Every record is stamped with the run id, process id, and runtime session id before the event's own fields.

    That answers both halves of this issue directly:

    • Ownership — each file is owned by exactly one process for its lifetime, keyed on run id (tracePath/trajectoryPath at src/trace.zig:43-49). There is no shared cursor to race on, so the startup-truncation and same-append-position failure modes are structurally gone rather than merely locked around.
    • Integrity — records stay pre-serialized and flushed whole, and now additionally carry run id + pid + session id, so a reader can attribute every line.

    Readers were updated to match: the archive reader scans all .graff/trajectories/*.jsonl files rather than assuming one (src/trace.zig:336).

    Reopen if a shared-file path still exists somewhere I have not spotted.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions