Repository navigation
Conversation
Three stdlib-only scripts, written to check another lane's claim rather than
to change behaviour. No production code is touched.
verify/identity_check.py fixed-seed simulate_ssa over 25 repository models plus
synthetic edge cases (dimerization, zero-rate plateaus, a
species on both sides of the fired reaction, a negative
rate constant forcing the non-monotone selection
fallback, a wide chain, a long ring); hashes every
.gdat/.cdat/.net artifact.
verify/compare_identity.py diffs two such trees per model, separating IDENTICAL,
DIFFERS, and INCONCLUSIVE.
verify/ab_bench.py interleaved A/B throughput harness: child CPU time, a
paired 2-event probe for overhead correction, per-round
order flip, and a .gdat hash guard that fails the run if
any arm or rep diverges.
Reproduces the check behind: perfEngine's 1.599x SSA claim re-measured at 2.061x
min / 2.044x median raw and 2.129x / 2.120x probe-corrected, arms non-overlapping,
against a pair differing in one TU pinned by blob SHA (57e504a vs 07baa1a);
22/26 models byte-identical with the risk-surface cases named, two inconclusive
because the BASELINE binary is nondeterministic there, one excluded on time;
exact integer conservation over 300,000 events as an absolute check the hash
cannot provide; and orchPerf's reported depOffset[idx+1] out-of-bounds write NOT
reproducible across 2711 .net files.
Member
Author
|
Superseded by reviewed PR #92, merged as 7107b6c. All three identity/A-B harness paths are carried as reviewed updated files, including fresh usable artifacts, complete sampling grids, comparable manifests and exact nonempty case coverage with regressions. Old trajectory outputs do not qualify under the repaired contract. PR92 head6d835c3 finished40successful/5skipped checks. Original author branch is preserved. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Harness only — no production code. Three stdlib-only scripts, written to check another lane's claim rather than to change behaviour.
What this reproduces
An independent re-measurement of
agent/perf-engine's SSA claim (PR #44), built from source in a separate worktree with a matched binary pair differing in one TU, pinned by blob SHA (57e504a= origin/main 6889fba vs07baa1a= aab2c0f;git diff --stat= 1 file, +9/-93).depOffset[idx+1]OOB write (orchPerf).netfiles, 792,960 reaction lines, zero out-of-range refsArms do not overlap (base max 910k < changed min 1.79M events/sec) across 9 interleaved reps with per-round order flip. Peak child RSS identical at 21,463,040 kB.
Limits, stated rather than rounded up
isTotalRatepath, both identical).Measurement context
R=48 runnable, loadavg band 117–147 (coarse band only). Three declared co-tenants during the timed window, not two: my A/B arms, perfEngine's NFsim baseline, and my own identity sweep — an omission I reported and corrected. Nothing load-bearing here is a configure, ctest, or build-system gate, which is why these results are unaffected by the duplicate-target defect or the FetchContent question.
The scripts
verify/identity_check.py— fixed-seedsimulate_ssaover 25 models plus synthetic edge cases targeting the change's risk surface: negative rate constant (non-monotone fallback), zero-rate plateaus, identical reactants, a species on both sides of the fired reaction, larger reaction counts, and saturating rates where propensity reads no species at all.verify/compare_identity.py— separates IDENTICAL / DIFFERS / INCONCLUSIVE.verify/ab_bench.py— interleaved A/B, child CPU time, paired 2-event probe, per-round order flip, and a.gdathash guard that fails the run if any arm or rep diverges.No Python package is imported by any of them, so the editable-install shadowing and the conftest
sys.pathfilter do not apply.🤖 Generated with Claude Code