You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Remove debug names/symbols from published release builds; preserve them separately for debugging/profiling. Merged: R2R metadata stripping defaults #134690, configurable wasm names/symbol-map sidecars #134808 (names still enabled by default).
Investigate whether we can avoid interpreter -> R2R calls for simple intrinsics (volatile memory operations, Unsafe, MemoryMarshal, bit rotations).
Analyze more slow microbenchmarks, identify their bottlenecks, and verify improvements under the release setup.
Release setup: composite R2R + PGO
Target composite R2R with profile-guided optimization (PGO) for the release configuration. Establish representative training workloads and validate the combination rather than assuming its gains.
Havit experiment (no PGO yet): coherent runtime cfe8a6c4, same benchmark CLI, 10 cold/5 warm samples per mode. Composite ran cleanly but startup was 542 versus 467 ms cold; 318 versus 275 ms warm, approximately 16% slower than per-assembly R2R, despite approximately 4.2x fewer cold transitions. Download grew 7.5%, with a 5.3 MiB owner image versus 2.1 MiB for the largest per-assembly module. A larger module's compilation/instantiation cost may outweigh fewer transitions; this explanation is not yet CPU-profiled. One walkthrough sample was 8.4% faster with composite, but needs repetition. This is not a measured composite + PGO benefit.
Browser composite publishing landed in #134618, which reports IndexOfMax<double>(3079) running the SIMD path at 4,040 ns/op composite versus 10,652,060 ns/op per-assembly R2R on an earlier PR revision. @lewing, which compilation/inlining changes explain this, and do other benchmarks show similar gains?
Startup performance and size
Local Havit at runtime 8970fe8 (Release, desktop macOS/arm64; 10 cold/5 warm samples) improved time to reach managed code from 1006/519 ms to 721/286 ms; download fell 12.6%. This combined PublishReadyToRunStripDebugInfo=true, PublishReadyToRunStripInliningInfo=true, and wasm name-section removal; it does not isolate names alone. The host was busy; these are same-session historical comparisons, not current-build baselines.
Performance improvements
R2R <-> interpreter transitions
Local Havit startup with coherent per-assembly R2R at runtime cfe8a6c4 recorded 313,286 cold / 310,353 warm transitions. Cold: 213,892 interpreter -> R2R and 99,394 R2R -> interpreter. The earlier September 30 SDK run, using mixed-vintage bits, recorded approximately 2.27 million. These are instrumented startup counts, not walkthrough counts or timings.
#134078 added wasm R2R generic helpers and is a likely contributor; the approximately 7.2x reduction spans multiple changes and has not been bisected.
Of 98,178 attributed reverse hits, 58.6% were ASP.NET Core Components, 16.6% URI parsing, and 14.2% virtual-dispatch/unboxing thunks. Forward attribution uses 160 samples; current targets include reflection/Type and Unsafe/Span primitives. Earlier samples also identified Volatile.Read, Volatile.ReadBarrier, Unsafe.As<T>, Unsafe.Add, MemoryMarshal.GetArrayDataReference<T>, and BitOperations.RotateLeft/RotateRight.
Should these small operations be expanded directly by the interpreter? Verify affected overloads and existing support first.
Cached interpreter dispatch already landed in #123815.
Microbenchmark investigation
Concrete gap: System.Collections.IterateForEach<Int32>.Span(Size: 512) (4. in the linked ranking). A coherent local run at runtime d90dbf43be15, performance 70a3326e, V8 15.4.80 (15 samples per mode):
Mono AOT
CoreCLR per-assembly R2R
CoreCLR composite R2R
2.724 ns
24.793 µs
187.952 ns
Per-assembly R2R lacks the closed int body and executes all 512 iterations interpreted. Composite emits that native body and removes the interpreter transition from the measured path: 131.9x faster. It still executes the loop; Mono/LLVM collapses it to a last-element load. The remaining 69x gap is native/native but different generated work. This demonstrates compiled-method coverage, not intrinsic inlining or a measured PGO benefit.
Runtime tests
Revert changes to SIMD tests once performance improves enough.
Revert changes to Regression_o_1 tests once performance improves enough.
Record revisions, engine/configuration, sample counts, and actual execution modes for each comparison.
Note
This issue update was prepared with GitHub Copilot assistance.
Composite mode is causing the uncompressed binary to get substantially larger, even though the compressed size is smaller than non-composite on my local measurements. What this means is that with composite mode we are compiling a LOT of nearly duplicate methods. This is likely causing the initial validation of the WebAssembly file to be a fair bit slower. I spent some time theorizing improvements to this last week, and will likely be trying out some of them this week.
Track .NET 12 WebAssembly/CoreCLR performance. Shared interpreter work: #122464.
Performance dashboards and investigation targets
The ranking uses build
20260913.1; refresh it after recent R2R improvements. Validate near-zero reference timings before prioritizing by ratio.Next steps
Release setup: composite R2R + PGO
Target composite R2R with profile-guided optimization (PGO) for the release configuration. Establish representative training workloads and validate the combination rather than assuming its gains.
Havit experiment (no PGO yet): coherent runtime
cfe8a6c4, same benchmark CLI, 10 cold/5 warm samples per mode. Composite ran cleanly but startup was 542 versus 467 ms cold; 318 versus 275 ms warm, approximately 16% slower than per-assembly R2R, despite approximately 4.2x fewer cold transitions. Download grew 7.5%, with a 5.3 MiB owner image versus 2.1 MiB for the largest per-assembly module. A larger module's compilation/instantiation cost may outweigh fewer transitions; this explanation is not yet CPU-profiled. One walkthrough sample was 8.4% faster with composite, but needs repetition. This is not a measured composite + PGO benefit.Browser composite publishing landed in #134618, which reports
IndexOfMax<double>(3079)running the SIMD path at 4,040 ns/op composite versus 10,652,060 ns/op per-assembly R2R on an earlier PR revision. @lewing, which compilation/inlining changes explain this, and do other benchmarks show similar gains?Startup performance and size
Local Havit at runtime
8970fe8(Release, desktop macOS/arm64; 10 cold/5 warm samples) improved time to reach managed code from 1006/519 ms to 721/286 ms; download fell 12.6%. This combinedPublishReadyToRunStripDebugInfo=true,PublishReadyToRunStripInliningInfo=true, and wasm name-section removal; it does not isolate names alone. The host was busy; these are same-session historical comparisons, not current-build baselines.Performance improvements
R2R <-> interpreter transitions
Local Havit startup with coherent per-assembly R2R at runtime
cfe8a6c4recorded 313,286 cold / 310,353 warm transitions. Cold: 213,892 interpreter -> R2R and 99,394 R2R -> interpreter. The earlier September 30 SDK run, using mixed-vintage bits, recorded approximately 2.27 million. These are instrumented startup counts, not walkthrough counts or timings.#134078 added wasm R2R generic helpers and is a likely contributor; the approximately 7.2x reduction spans multiple changes and has not been bisected.
Of 98,178 attributed reverse hits, 58.6% were ASP.NET Core Components, 16.6% URI parsing, and 14.2% virtual-dispatch/unboxing thunks. Forward attribution uses 160 samples; current targets include reflection/Type and Unsafe/Span primitives. Earlier samples also identified
Volatile.Read,Volatile.ReadBarrier,Unsafe.As<T>,Unsafe.Add,MemoryMarshal.GetArrayDataReference<T>, andBitOperations.RotateLeft/RotateRight.Should these small operations be expanded directly by the interpreter? Verify affected overloads and existing support first.
Cached interpreter dispatch already landed in #123815.
Microbenchmark investigation
Concrete gap:
System.Collections.IterateForEach<Int32>.Span(Size: 512)(4. in the linked ranking). A coherent local run at runtimed90dbf43be15, performance70a3326e, V815.4.80(15 samples per mode):Per-assembly R2R lacks the closed
intbody and executes all 512 iterations interpreted. Composite emits that native body and removes the interpreter transition from the measured path: 131.9x faster. It still executes the loop; Mono/LLVM collapses it to a last-element load. The remaining 69x gap is native/native but different generated work. This demonstrates compiled-method coverage, not intrinsic inlining or a measured PGO benefit.Runtime tests
Regression_o_1tests once performance improves enough.Record revisions, engine/configuration, sample counts, and actual execution modes for each comparison.
Note
This issue update was prepared with GitHub Copilot assistance.