Skip to content

WIP(netgen): network-generation performance work — benchmark harness + measured candidates (evidence pending) - #42

Closed
akutuva21 wants to merge 2 commits into
mainfrom
agent/perf-netgen
Closed

akutuva21 wants to merge 2 commits into
mainfrom
agent/perf-netgen

Conversation

@akutuva21

Copy link
Copy Markdown
Member

Status

WIP / evidence-pending. Benchmark harness and fixtures are committed; the first measured performance candidate (copy-where-move-was-available in cpp/ast/RefineRule.cpp) is in flight. Numbers will follow in this body as each change lands with its own A/B table.

What changed so far (bench only, no production changes yet)

  • bench/netgen_bench.py — reproducible network-generation benchmark:
    • 11 generate-only fixtures (bench/models/*_gen.bngl), regenerated deterministically by bench/make_fixtures.py from models/;
    • heavy preset (egfr_net, fceri_ji, e6, e7) for A/B loops, --full for the 11-fixture final gate;
    • arms interleaved per round (--binary A --binary B), per-rep times, min/median/mean/stdev, sorted totals, noise warning when median−min > 25% of min;
    • measurement context printed: loadavg as a coarse band + runnable-process count at start/end (loadavg decimals are untrustworthy on this host — memWatch calibration), per-fixture peak RSS from wait4 rusage;
    • correctness gate: raw .net byte-identity across all reps and binaries, with one documented exemption (below).
  • bench/models/*.bngl — 11 fixtures, each a source model truncated after its last section terminator with a single generate_network action reusing the model's own arguments.

tlbr exemption (pre-existing defect, not mine, owned elsewhere)

bench/models/tlbr_gen.bngl is exempt from the raw-hash gate because the BASELINE binary itself emits two site-numbering variants of the same network across runs:

for i in $(seq 1 60); do build/cpp/bng_cpp bench/models/tlbr_gen.bngl >/dev/null 2>&1; shasum -a 256 bench/models/tlbr_gen.net; done | sort | uniq -c
-> 46 runs hash 9630a77b..., 14 runs hash 2aa4b87b...   (this binary)
independent reproduction on another binary of the same base: 36/4 over 40 runs; shared main binary: 28/12 over 40 runs

Diff between the two variants (reactions block identical; species lines differ in which site carries the partner bond, e.g. line 21: R(l!2,l!4) vs R(l!3,l!4)). Mechanism is a hypothesis (address-dependent embedding selection → different product-graph construction order), reported, not fixed here. tlbr still runs in the benchmark for timing diversity and must satisfy a weak structural gate (identical species-line count, identical reactions block, identical groups block across all reps and binaries). The other 10 fixtures require byte-identical .net.

Methodology (standing rules this PR follows)

  • Benchmark before/after, one change at a time, each with its own interleaved A/B table.
  • Numbers carry: loadavg band, exact concurrent process list, interleaving statement, peak RSS (/usr/bin/time -l / wait4), per Main's measurement-context rule.
  • Host cross-session drift floor ~11%: claims below it are reported as within-noise unless re-measured; allocation counts (deterministic, verified by an independent harness) are the primary metric for allocation-class changes.
  • Independent verification: swarmMemory links archived baseline/candidate static libs with its in-process allocation counter and reports process-TOTAL deltas with a relocation guard.

Verification (will be quoted here per change)

  • python3 bench/netgen_bench.py --reps 5 --binary /tmp/netgen_bench/bng_cpp_baseline --binary build/cpp/bng_cpp
    ctest --test-dir build --output-on-failure
  • Baseline captured before any production edit: commit 90ba29e/0160334 tree, heavy preset, see body updates below.

@akutuva21

Copy link
Copy Markdown
Member Author

Superseded by reviewed PR #92, merged as 7107b6c. Exact residual audit confirms all 13 benchmark paths are carried (2 identical, 11 updated); no production optimization was omitted. PR92 head6d835c3 finished40successful/5skipped checks. This disposition carries benchmark scaffolding and bounded evidence, not a general speedup claim. Original author branch is preserved.

@akutuva21 akutuva21 closed this Oct 4, 2026
@akutuva21
akutuva21 deleted the agent/perf-netgen branch October 4, 2026 02:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant