Skip to content

Perf: retain checkpoint-safe object allocation reserves #125

Description

@chrisbbreuer

Parent: #97
Related: zig-utils/zig-gc#25, zig-utils/zig-gc#26

Evidence

The post-#120 exact shared object_churn profile resolves the hottest allocation wait to zig-gc metadata publication. zig-gc#26 removes the failed-CAS write storm and improves the exact eight-lane median 11.8%, but the VM still publishes a batch about every 17 objects because each quickened loop invocation stops at the 1,024-step checkpoint.

zig-gc#25 proved that shortening a 17-cell critical section alone is not general: its O(1) list splice regressed the lower shared lanes. The remaining lever belongs in zig-js: retain a bounded initialized-cell reserve as an explicit interpreter GC root so one larger publication can serve several checkpoint-bounded invocations safely.

Scope

  • Add a bounded per-interpreter reserve used only by the guarded fixed-shape object-allocation quick path in parallel mode.
  • Default-initialize every published cell before it enters the reserve and precisely trace only its unused suffix at every collection/root-publication path.
  • Consume cells in allocation order, preserve short-prefix/OOM behavior, and never carry an unpublished or uninitialized cell across a safepoint.
  • Add telemetry proving the existing workload publishes materially fewer batches, plus focused collection/root-survival coverage.
  • Benchmark direct, shared 1/2/4/8, and independent-context behavior with exact checksums and bounded footprint; reject the change if lower lanes regress.

Acceptance

  • Preserve exact object semantics, observable step/checkpoint order, GC safety, and allocation-failure behavior.
  • Reduce shared object-batch publications by at least 8x on the focused allocation loop.
  • Materially improve shared 4/8-lane wall time without regressing direct or independent-context results.
  • Pass focused semantic/GC tests, full units, and TSan before publication.

Activity

  1. chrisbbreuer commented on Jul 16, 2026

    @chrisbbreuer
    MemberAuthor

    Rejected after exact-checksum ReleaseFast A/B; no source changes were retained. A GC-traced per-interpreter reserve was implemented and its focused shared-thread unit test passed 2/2 with zero leaks, but larger publication bursts worsened the shared convoy.\n\nSeven-sample shared object_churn medians (100 jobs):\n- 256 cells: 1 lane 124.144 ms vs parent 135.661 ms (-8.5%); 2 lanes 241.087 vs 233.034 ms (+3.5%); 4 lanes 470.105 vs 432.262 ms (+8.8%); 8 lanes 1,184.045 vs 1,052.498 ms (+12.5%).\n- 64 cells: 4 lanes 669.659 vs parent 477.814 ms (+40.1%); 8 lanes 1,221.497 vs 1,077.187 ms (+13.4%).\n\nEvery checksum matched. The source was restored exactly and the focused build artifacts are temporary only. Together with zig-gc#25, this rules out both sides of the coarse-batching idea: O(1) splicing at ~17 cells and carrying 64/256 initialized roots across checkpoints. The next #97 slice must remove or shard shared allocation metadata, not merely change batch duration/frequency.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions