Skip to content

[wasm][coreclr] Performance tracking issue #124218

Description

@radekdoulik

Track .NET 12 WebAssembly/CoreCLR performance. Shared interpreter work: #122464.

Performance dashboards and investigation targets

The ranking uses build 20260913.1; refresh it after recent R2R improvements. Validate near-zero reference timings before prioritizing by ratio.

Next steps

  • Set up and validate composite R2R + PGO for release builds. Merged support: composite #134618, SDK output naming dotnet/sdk#56395, interpreter PGO collection #132721, wasm MIBC consumption without embedded PGO data #134093.
  • Enable the composite R2R microbenchmark lane: dotnet/performance#5324.
  • Remove debug names/symbols from published release builds; preserve them separately for debugging/profiling. Merged: R2R metadata stripping defaults #134690, configurable wasm names/symbol-map sidecars #134808 (names still enabled by default).
  • Investigate whether we can avoid interpreter -> R2R calls for simple intrinsics (volatile memory operations, Unsafe, MemoryMarshal, bit rotations).
  • Analyze more slow microbenchmarks, identify their bottlenecks, and verify improvements under the release setup.

Release setup: composite R2R + PGO

Target composite R2R with profile-guided optimization (PGO) for the release configuration. Establish representative training workloads and validate the combination rather than assuming its gains.

Havit experiment (no PGO yet): coherent runtime cfe8a6c4, same benchmark CLI, 10 cold/5 warm samples per mode. Composite ran cleanly but startup was 542 versus 467 ms cold; 318 versus 275 ms warm, approximately 16% slower than per-assembly R2R, despite approximately 4.2x fewer cold transitions. Download grew 7.5%, with a 5.3 MiB owner image versus 2.1 MiB for the largest per-assembly module. A larger module's compilation/instantiation cost may outweigh fewer transitions; this explanation is not yet CPU-profiled. One walkthrough sample was 8.4% faster with composite, but needs repetition. This is not a measured composite + PGO benefit.

Browser composite publishing landed in #134618, which reports IndexOfMax<double>(3079) running the SIMD path at 4,040 ns/op composite versus 10,652,060 ns/op per-assembly R2R on an earlier PR revision. @lewing, which compilation/inlining changes explain this, and do other benchmarks show similar gains?

Startup performance and size

Local Havit at runtime 8970fe8 (Release, desktop macOS/arm64; 10 cold/5 warm samples) improved time to reach managed code from 1006/519 ms to 721/286 ms; download fell 12.6%. This combined PublishReadyToRunStripDebugInfo=true, PublishReadyToRunStripInliningInfo=true, and wasm name-section removal; it does not isolate names alone. The host was busy; these are same-session historical comparisons, not current-build baselines.

Performance improvements

R2R <-> interpreter transitions

Local Havit startup with coherent per-assembly R2R at runtime cfe8a6c4 recorded 313,286 cold / 310,353 warm transitions. Cold: 213,892 interpreter -> R2R and 99,394 R2R -> interpreter. The earlier September 30 SDK run, using mixed-vintage bits, recorded approximately 2.27 million. These are instrumented startup counts, not walkthrough counts or timings.

#134078 added wasm R2R generic helpers and is a likely contributor; the approximately 7.2x reduction spans multiple changes and has not been bisected.

Of 98,178 attributed reverse hits, 58.6% were ASP.NET Core Components, 16.6% URI parsing, and 14.2% virtual-dispatch/unboxing thunks. Forward attribution uses 160 samples; current targets include reflection/Type and Unsafe/Span primitives. Earlier samples also identified Volatile.Read, Volatile.ReadBarrier, Unsafe.As<T>, Unsafe.Add, MemoryMarshal.GetArrayDataReference<T>, and BitOperations.RotateLeft/RotateRight.

Should these small operations be expanded directly by the interpreter? Verify affected overloads and existing support first.

Cached interpreter dispatch already landed in #123815.

Microbenchmark investigation

Concrete gap: System.Collections.IterateForEach<Int32>.Span(Size: 512) (4. in the linked ranking). A coherent local run at runtime d90dbf43be15, performance 70a3326e, V8 15.4.80 (15 samples per mode):

Mono AOT CoreCLR per-assembly R2R CoreCLR composite R2R
2.724 ns 24.793 µs 187.952 ns

Per-assembly R2R lacks the closed int body and executes all 512 iterations interpreted. Composite emits that native body and removes the interpreter transition from the measured path: 131.9x faster. It still executes the loop; Mono/LLVM collapses it to a last-element load. The remaining 69x gap is native/native but different generated work. This demonstrates compiled-method coverage, not intrinsic inlining or a measured PGO benefit.

Runtime tests

  • Revert changes to SIMD tests once performance improves enough.
  • Revert changes to Regression_o_1 tests once performance improves enough.

Record revisions, engine/configuration, sample counts, and actual execution modes for each comparison.

Note

This issue update was prepared with GitHub Copilot assistance.

Activity

  1. added this to the Future milestone on Feb 10, 2026
  2. dotnet-policy-service commented on Feb 10, 2026

    @dotnet-policy-service
    Contributor

    Tagging subscribers to 'arch-wasm': @lewing, @pavelsavara
    See info in area-owners.md if you want to be subscribed.

  3. dotnet-policy-service commented on Feb 10, 2026

    @dotnet-policy-service
    Contributor

    Tagging subscribers to this area: @dotnet/interop-contrib
    See info in area-owners.md if you want to be subscribed.

  4. added
    area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI
    and removed on Oct 4, 2026
  5. dotnet-policy-service commented on Oct 4, 2026

    @dotnet-policy-service
    Contributor

    Tagging subscribers to this area: @JulieLeeMSFT, @jakobbotsch
    See info in area-owners.md if you want to be subscribed.

  6. davidwrighton commented on Oct 5, 2026

    @davidwrighton
    Member

    Composite mode is causing the uncompressed binary to get substantially larger, even though the compressed size is smaller than non-composite on my local measurements. What this means is that with composite mode we are compiling a LOT of nearly duplicate methods. This is likely causing the initial validation of the WebAssembly file to be a fair bit slower. I spent some time theorizing improvements to this last week, and will likely be trying out some of them this week.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    arch-wasmWebAssembly architecturearea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI

    Type

    No type

    Projects

    • Status
      No status

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions