Skip to content

verify: independent re-measurement harness for the SSA throughput claim - #78

Closed
akutuva21 wants to merge 1 commit into
mainfrom
verify/harness-on-main
Closed

akutuva21 wants to merge 1 commit into
mainfrom
verify/harness-on-main

Conversation

@akutuva21

Copy link
Copy Markdown
Member

Harness only — no production code. Three stdlib-only scripts, written to check another lane's claim rather than to change behaviour.

What this reproduces

An independent re-measurement of agent/perf-engine's SSA claim (PR #44), built from source in a separate worktree with a matched binary pair differing in one TU, pinned by blob SHA (57e504a = origin/main 6889fba vs 07baa1a = aab2c0f; git diff --stat = 1 file, +9/-93).

claimed measured verdict
throughput (species=150, R=450, 2M-event cap) 1.599x 2.061x min / 2.044x median raw; 2.129x / 2.120x probe-corrected reproduced, conservative
bit-identical trajectories yes 22/26 models byte-identical reproduced
trajectory physics correct — exact integer conservation, drift 0.0 over 300,000 events reproduced
depOffset[idx+1] OOB write (orchPerf) flagged 2711 .net files, 792,960 reaction lines, zero out-of-range refs not reproduced

Arms do not overlap (base max 910k < changed min 1.79M events/sec) across 9 interleaved reps with per-round order flip. Peak child RSS identical at 21,463,040 kB.

Limits, stated rather than rounded up

  • Measured at one operating point (R=450). This does not replace perfEngine's scaling table and says nothing at R=45 or R=1350.
  • Identity: 2 models INCONCLUSIVE — three identical runs of the baseline binary produce different hashes, so no comparison can adjudicate them. Reported as neither win nor defect. 1 excluded (exceeded a 240s cap, no artifacts in either arm). 1 not-a-case (stale dir, replaced by test_MM/test_sat covering the same isTotalRate path, both identical).
  • The identity gate is a test of the trajectory, not the engine, and has a half-life: if the SSA seed derivation changes (as fix(batch-ssa): restore the missing batchTrajectorySeed declaration — main does not compile without it #69 did), these hashes become vacuous rather than false. The conservation check is the absolute instrument and does not share that bound.
  • perfBatch's PR perf(batch-ssa): compile the model once for the whole CPU pool #58 was partially verified: −1.80% vs −2.0% claimed on the small-batch config; the headline −36.4% and RSS halving were not measured here (three attempts exceeded 3 minutes each; killed rather than report a partial figure).

Measurement context

R=48 runnable, loadavg band 117–147 (coarse band only). Three declared co-tenants during the timed window, not two: my A/B arms, perfEngine's NFsim baseline, and my own identity sweep — an omission I reported and corrected. Nothing load-bearing here is a configure, ctest, or build-system gate, which is why these results are unaffected by the duplicate-target defect or the FetchContent question.

The scripts

  • verify/identity_check.py — fixed-seed simulate_ssa over 25 models plus synthetic edge cases targeting the change's risk surface: negative rate constant (non-monotone fallback), zero-rate plateaus, identical reactants, a species on both sides of the fired reaction, larger reaction counts, and saturating rates where propensity reads no species at all.
  • verify/compare_identity.py — separates IDENTICAL / DIFFERS / INCONCLUSIVE.
  • verify/ab_bench.py — interleaved A/B, child CPU time, paired 2-event probe, per-round order flip, and a .gdat hash guard that fails the run if any arm or rep diverges.

No Python package is imported by any of them, so the editable-install shadowing and the conftest sys.path filter do not apply.

🤖 Generated with Claude Code

Three stdlib-only scripts, written to check another lane's claim rather than
to change behaviour. No production code is touched.

verify/identity_check.py   fixed-seed simulate_ssa over 25 repository models plus
                          synthetic edge cases (dimerization, zero-rate plateaus, a
                          species on both sides of the fired reaction, a negative
                          rate constant forcing the non-monotone selection
                          fallback, a wide chain, a long ring); hashes every
                          .gdat/.cdat/.net artifact.
verify/compare_identity.py diffs two such trees per model, separating IDENTICAL,
                          DIFFERS, and INCONCLUSIVE.
verify/ab_bench.py        interleaved A/B throughput harness: child CPU time, a
                          paired 2-event probe for overhead correction, per-round
                          order flip, and a .gdat hash guard that fails the run if
                          any arm or rep diverges.

Reproduces the check behind: perfEngine's 1.599x SSA claim re-measured at 2.061x
min / 2.044x median raw and 2.129x / 2.120x probe-corrected, arms non-overlapping,
against a pair differing in one TU pinned by blob SHA (57e504a vs 07baa1a);
22/26 models byte-identical with the risk-surface cases named, two inconclusive
because the BASELINE binary is nondeterministic there, one excluded on time;
exact integer conservation over 300,000 events as an absolute check the hash
cannot provide; and orchPerf's reported depOffset[idx+1] out-of-bounds write NOT
reproducible across 2711 .net files.
@akutuva21

Copy link
Copy Markdown
Member Author

Superseded by reviewed PR #92, merged as 7107b6c. All three identity/A-B harness paths are carried as reviewed updated files, including fresh usable artifacts, complete sampling grids, comparable manifests and exact nonempty case coverage with regressions. Old trajectory outputs do not qualify under the repaired contract. PR92 head6d835c3 finished40successful/5skipped checks. Original author branch is preserved.

@akutuva21 akutuva21 closed this Oct 4, 2026
@akutuva21
akutuva21 deleted the verify/harness-on-main branch October 4, 2026 02:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant