Repository navigation
docs(perf): correct the +0.024% figure published in #38 - #52
Merged
Merged
Conversation
akutuva21
force-pushed
the
swarm/compiler-docfix
branch
3 times, most recently
from
October 1, 2026 02:56
3c7859d to
0c0b5d0
Compare
PR #38 merged at 25da219, before this correction landed, so the record on main currently states a resolvable-looking delta that is not resolvable. Re-analysing both independent instruction-count datasets: instr.txt A_med 5,039,373,445 B_med 5,037,690,571 -0.033% B FASTER instr2.txt A_med 5,037,345,446 B_med 5,038,543,790 +0.024% B SLOWER The sign flips between datasets, and the between-arm delta (~0.03%) is smaller than the within-arm full range (0.061-0.099% of median). The correct reading is 'no measurable difference in either direction'. Same class of error as the original 23.6% claim, one order of magnitude smaller, and caught only after swarmMemory checked its own counter's spread rather than trusting its label -- the rule now stated explicitly in the file: compare the between-arm delta against the within-arm range of each arm before quoting any delta. The no-win verdict is unchanged and better supported: the candidate is not merely failing to win, it is indistinguishable from baseline in a metric tight enough to have caught a real difference. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The document told the next operator that writeOutputFiles is 'where a codegen-shaped win would actually land'. Two things learned since make that actively misleading rather than merely unhelpful: 1. cpp/engine/OdeIntegrator.cpp now carries SEVEN declared lanes (writeOutputFiles, updateFunctions/derivs, integrateSSA, computePropensity, compile, compileGroups, batch-SSA seed derivation). Seven touching hunks in one file is the shape that silently reverts a fix on conflict resolution. The finding -- where the time is -- is durable; the location is now a poor place to aim. 2. The win that landed there was not the shape I predicted. swarmSerial measured 2.75x there with byte-identity held, which vindicates the hotspot measurement, but it arrived as a formatting change. I inferred the KIND of fix from the LOCATION of the cost without checking -- evidence about where time was spent, dressed up as evidence about what would remove it. Same error class as the rest of this file. Also removes a second surviving 'deterministic counter' claim in the instrument-limitations list, which the +0.024% correction superseded elsewhere in the file but missed here. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
swarmCache's split is sharper than what I had: a hash guard is valid when the FIXTURE is deterministic, not when the change is value-preserving. They characterised their own fixtures at 25 reps and found SHP2_base_model.bngl producing 6 distinct .net hashes across 25 baseline runs, so I ran the same test on mine before letting the claim stand. Same binary, two runs, 4000002 rows each: a27182ecb5f5451c08bc45677d3be6a36aefffb864a74b2ec96df3002bedd05b (both) So the fixture is deterministic and the two-binary comparison is real evidence. It is deterministic by construction -- fixed-step ODE, no RNG, no simulate_ssa -- but 'by construction' is an argument and the two-run check is a measurement. I had implied the precondition rather than shown it. Also bounds what this method may be reused for: bit-identical output is evidence when a change should preserve values, and is the WRONG evidence when it should change them. Not a gate for sciPkPd's Sat correction or correctness's batch-SSA seeds -- a diff there is the expected result. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
swarmCache named the trap precisely: '3 distinct hashes' from a pre-fix binary and from a post-fix binary mean opposite things, and the reader cannot tell which from a hash count. It applies to the hashes I published. Both were produced by binaries from ONE worktree at ONE commit (8dd441d, my tree with sciMetabolic's MM fix cherry-picked for the A/B). They are evidence internal to that comparison only. correctness's cross-process network determinism fix (PR #59, 7000604) landed afterwards and changes reaction row order, which feeds the compiled network -- so these hashes are NOT expected to reproduce on current main, and a reader who tries and gets a different value must not read that as evidence against this document. Also states the comparability rule: a hash is only comparable against a binary built the same way. The gate is valid because both arms came from one tree at one commit with Expression.cpp the only difference. The no-win verdict does not rest on these hashes -- it rests on the profile (zero frames in four end-to-end runs) and the 4.89% ceiling, neither of which depends on any artifact reproducing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
akutuva21
force-pushed
the
swarm/compiler-docfix
branch
from
October 1, 2026 02:57
0c0b5d0 to
20d3c0a
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What is wrong on main
#38 merged at
25da219, before I found the error. The merged document reportsthe instruction counter as settling the question at
+0.024%, i.e. thecandidate slower, phrased as though the counter had resolved a difference.
It had not. Re-analysing the two independent datasets I collected:
The sign flips between datasets, and the between-arm delta (~0.03%) is
smaller than the within-arm full range (0.061%–0.099% of median). The honest
reading is no measurable difference in either direction.
Same class of error as the original retracted 23.6% claim, one order of
magnitude smaller. It survived because I trusted my own counter's label instead
of measuring its spread — the check
swarmMemoryapplied to its own allocationcounters after finding a 13.69% full range under a "deterministic, min==max"
label. Mine had a 0.06–0.10% range and I still quoted 0.024% against it.
What this PR changes
states what actually mattered: the counter contradicted wall clock's
+4.78%, and that contradiction is what killed the candidate.
datasets, delta below within-arm noise, replacing the false precision.
within-arm range of each arm before quoting any delta. Tightness does not
license it — a tight-looking counter makes you more confident, not less, which
is exactly when the check gets skipped.
adjudicate a 23.6% claim, not a 0.03% one.
Second correction in this PR: the
writeOutputFilespointer is stale adviceThe document also told the next operator that
writeOutputFilesis "where acodegen-shaped win would actually land on this platform". That is now actively
misleading rather than merely unhelpful:
cpp/engine/OdeIntegrator.cppnow carries seven declared lanes —writeOutputFiles,updateFunctions/derivs,integrateSSA,computePropensity,compile,compileGroups, and the batch-SSA seedderivation. Seven touching hunks in one file is the shape that silently
reverts a fix when someone resolves a conflict by picking a side. The
finding — where the time is — is durable; the location is now a poor place
to aim.
swarmSerialmeasured 2.75x there (0.11s → 0.04s, byte-identity held across30 artifacts), which vindicates the hotspot measurement — but it arrived as a
formatting change. I inferred the kind of fix from the location of the cost
without checking, which is the same error class as the rest of this file:
evidence about where time was spent, dressed up as evidence about what would
remove it.
The transferable form, now stated in the document: a profile tells you where,
and nothing about what kind of change will help there.
This also removes a second surviving "deterministic counter" claim in the
instrument-limitations list that the
+0.024%correction superseded elsewherein the file but missed here.
The verdict is unchanged, and slightly better supported
The no-win stands. The candidate does not merely fail to win; it is
indistinguishable from baseline in a metric tight enough to have caught a real
difference. That is a stronger claim than "+0.024% slower", and it is the one I
would rather have defended.
Files:
docs/PERF_EXPRESSION_CODEGEN_NO_WIN.mdonly. No production code.🤖 Generated with Claude Code