Skip to content

Perf: publish quiescent owned batches without the allocation lock #27

Description

@chrisbbreuer

Consumer parent: zig-utils/zig-js#97
Related: #25, #26, zig-utils/zig-js#125

Evidence

The post-#26 shared-heap profile shows metadata-lock handoff still dominates fixed-shape allocation. Two controlled alternatives have now failed:

Changing exclusive-section length or frequency therefore cannot solve the convoy. Owned-slab heaps already classify managed pointers without the global list or payload map; while no mark is active, allocation only needs to make new headers discoverable for the next collection.

Proposed slice

  • For parallel bindings that guarantee all cells use owned storage, initialize a non-marking batch privately and publish it onto an atomic pending-header stack without taking alloc_lock.
  • Use an active-publisher counter plus the existing marking transition as a gate: a collector arms marking under alloc_lock, waits for pre-arm publishers to leave, folds the complete pending stack into all, then whitens/traces. A publisher that observes the transition falls back to the existing locked born-grey path.
  • Aggregate live/byte/young accounting atomically once per fast batch. Collection threshold reads remain atomic in parallel mode; collection and teardown consume exact quiescent values.
  • Retain the proven locked path for marking, concurrent born-cell handoff, non-owned storage, OOM prefixes, and single-mutator heaps.

Acceptance

  • Prove no cell can publish between the collector's final pending drain and whiten/barrier arming.
  • Preserve newest-first nursery prefix order, exact accounting, short-prefix/OOM ordering, and concurrent born-grey semantics.
  • Add multi-mutator overlap tests that force publication/mark transitions and verify list/header/counter integrity under TSan.
  • Pass the full zig-gc unit and TSan workflow.
  • Exact-parent downstream A/B materially improves zig-js shared 4/8-lane object churn without regressing 1/2 lanes, direct, or independent contexts.

Activity

  1. chrisbbreuer commented on Jul 16, 2026

    @chrisbbreuer
    MemberAuthor

    Rejected after complete implementation and downstream A/B; the dependency worktree is restored exactly and no commit was made. The active-publisher gate, pending atomic chain, aggregate counters, forced allocation/mark transition test, all 39 normal tests, and the full TSan test gate passed. Every zig-js checksum matched.\n\nSeven-sample ReleaseFast shared object_churn medians (100 jobs):\n- 1 lane: lock-free 132.207 ms vs parent 136.527 ms (-3.2%)\n- 2 lanes: 235.037 vs 237.682 ms (-1.1%)\n- 4 lanes: 579.181 vs 495.106 ms (+17.0%)\n- 8 lanes: 1,159.067 vs 1,029.700 ms (+12.6%)\n\nThe atomic pending-head and aggregate-counter cache lines become a new shared hotspot at 4/8 lanes. Together with #25 and zig-js#125, this rules out shortening the lock, lengthening batches, and replacing the lock with one global CAS lane. The remaining #97 work needs genuinely sharded/per-mutator metadata or less metadata per allocation, with collection consuming shards only at an existing rendezvous.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions