Skip to content

perf: preserve concrete codec types in generated code - #727

Merged
SunSi12138 merged 65 commits into
devfrom
issue-726-concrete-codec-types
Sep 22, 2026
Merged

SunSi12138 merged 65 commits into
devfrom
issue-726-concrete-codec-types

Conversation

@SunSi12138

Copy link
Copy Markdown
Owner

Implements #726.

This change preserves compile-time concrete Codec types in generated code when the final binding is statically fixed and legally nameable.

Key points:

  • local generated native Codecs and frozen custom Codecs use concrete field types;
  • authoritative instances still resolve through IRpcCodecProvider and cast once during generated object construction;
  • Adapter, referenced generated, and built-in/internal Codecs retain IRpcCodec fallback;
  • Stub response encode no longer re-erases a concrete Codec through a generic IRpcCodec helper;
  • streaming runtime boundaries remain unchanged;
  • generated DTO, Union, Collection, request, Proxy, and Stub dependencies are covered by regression tests.

Base: dev

Copy link
Copy Markdown
Owner Author

Performance follow-up after the custom-struct fix and same-binary controls:

  • The earlier NativeAOT nested request-deserialize +4.5% result did not reproduce. On head f13c210, three alternating same-machine A/B rounds measured candidate -3.82% wall / -3.81% CPU vs dev for the same case.
  • More importantly, the Small DTO decode control (whose generated Deserialize instruction sequence is unchanged) also moved by about -5.5% to -6.2% across the two NativeAOT images. Normalized disassembly is instruction-for-instruction identical (226 instructions; same 0x34b function size), so cross-image code/link layout is a material confound and the previous +4.5% must not be treated as a product regression.
  • Same-binary NativeAOT interface-entry -> concrete-entry deltas are modest and directionally positive:
    • Small: dev -1.35%, candidate -2.46%
    • List: dev -1.26%, candidate -0.65%
    • Nested outer entry: ~0% (dev +0.23%, candidate -0.24%)
      These isolate call-site dispatch without comparing two separately linked images.
  • Candidate still shows direct child calls in both JIT and NativeAOT disassembly; dev uses IRpcCodec interface dispatch cells.
  • Allocation is unchanged for the decode controls.
  • Current whole-RPC NativeAOT nested unary remains essentially neutral on wall time (-0.72%) with CPU noise around +1.3%; the evidence does not support claiming a large end-to-end throughput gain.
  • Response serialize remains the clearest signal in this run: PGO ON -3.14%, NativeAOT -5.01%.
  • Wire SHA-256, CodecHash, MethodId, ContractId, fingerprints and RpcAssemblyHash remain identical.

Run: https://github.com/SunSi12138/SharpLink/actions/runs/35681317541

Interpretation: the concrete-dispatch codegen transition is real, but absolute dev-vs-candidate NativeAOT microbench deltas are sensitive to image layout. Same-binary controls are the safer evidence for dispatch cost.

@SunSi12138
SunSi12138 marked this pull request as draft September 22, 2026 03:05
@SunSi12138
SunSi12138 marked this pull request as ready for review September 22, 2026 03:05

Copy link
Copy Markdown
Owner Author

Added a #720-style isolated dispatch microkernel on head 4daa536.

The timed sample calls one dedicated tight loop once; the loop itself performs 2,000,000 Codec calls in normal evidence runs. There is no per-item delegate/RPC/transport overhead. The probe is a sealed class implementing IRpcCodec; Serialize/Deserialize are NoInlining and do only minimal state work so the measurement isolates interface indirect call vs concrete sealed direct call.

Three alternating same-machine rounds:

mode operation interface ns/call concrete ns/call concrete delta
JIT + Dynamic PGO ON Serialize ~2.62 ~2.63 ~0%
JIT + Dynamic PGO ON Deserialize ~4.8-5.1 ~4.8-4.9 noise / ~0%
JIT + Dynamic PGO OFF Serialize ~3.27 ~2.62 ~-20%
JIT + Dynamic PGO OFF Deserialize ~5.8-6.2 ~4.9-5.0 ~-14% to -20%
NativeAOT Serialize ~2.31-2.32 ~2.32 ~0%
NativeAOT Deserialize ~2.87-3.18 ~2.04-2.31 ~-27% to -29%

Alloc/call is 0 in all probe rows.

Codegen confirms the intended shapes:

  • JIT interface loop: IRpcCodec.Serialize/Deserialize indirect call; concrete loop: direct DispatchProbeCodec call.
  • NativeAOT interface loop: __InterfaceDispatchCell call; concrete loop: direct relative call.
  • Example JIT loop body size: Serialize interface 93 B vs concrete 75 B; Deserialize interface 94 B vs concrete 74 B.

Interpretation matches #720's earlier signal: Dynamic PGO can erase most of the class-interface dispatch cost, while PGO-disabled JIT exposes a clear ~15-20% microkernel dispatch tax. NativeAOT also has a meaningful direct-call signal for this Deserialize shape. This is intentionally an upper/isolation measurement, not an end-to-end throughput claim; the generated DTO/RPC measurements remain the realistic-effect layer.

Run: https://github.com/SunSi12138/SharpLink/actions/runs/35682751219

Copy link
Copy Markdown
Owner Author

NativeAOT dispatch follow-up is now narrowed down with return-dependency and same-method layout controls (head ba849b9, run 35684919723).

What changed in the probe:

  • Serialize and Deserialize now have symmetric minimal no-inline bodies (one state update; Deserialize only additionally returns the state).
  • Added Deserialize cases where the return value is intentionally unused.
  • Added same-method A/B controls: interface and concrete loops live in the same method and use the same Codec instance; a loop-external boolean selects the path. A/B invert source branch order to expose layout sensitivity.

Latest same-machine result:

mode control interface ns/call concrete ns/call direct delta
JIT + Dynamic PGO ON Serialize 3.279 3.248 -0.94%
JIT + Dynamic PGO ON Deserialize (used) 3.354 3.257 -2.88%
JIT + Dynamic PGO OFF Serialize 4.110 3.278 -20.32%
JIT + Dynamic PGO OFF Deserialize (used) 4.133 3.274 -20.80%
NativeAOT Serialize 2.808 1.561 -44.39%
NativeAOT Deserialize (used) 2.810 1.559 -44.55%
NativeAOT Deserialize (unused) 2.495 1.560 -37.50%

Same-method layout controls:

  • PGO ON: direct is essentially neutral for Serialize (~-0.5%) and ~-2.8% for Deserialize.
  • PGO OFF: both A/B layouts consistently show ~20-22% direct-call improvement.
  • NativeAOT: every A/B layout still favors concrete direct calls. Magnitude varies with layout (~22-45% in this run), so the absolute percentage is layout-sensitive, but the direction is robust.

This rules out the earlier hypotheses that the NativeAOT signal was primarily caused by ReadOnlySequence.Length, Deserialize return-value consumption, or simply having separate benchmark methods. The structural difference remains the call shape:

  • interface path: __InterfaceDispatchCell / indirect call
  • concrete path: direct relative call

The right interpretation is therefore:

  1. Dynamic PGO largely erases sealed-class interface dispatch cost.
  2. Without PGO, JIT pays a repeatable ~20% tax in this intentionally tiny call microkernel.
  3. NativeAOT has a real per-call interface-dispatch cost, but the exact percentage is strongly affected by native code layout; do not extrapolate the ~25-45% microkernel number to RPC throughput.
  4. The realistic generated-code layer remains much smaller: latest NativeAOT nested response serialize is ~-4.8%; whole nested unary is effectively neutral; cross-image decode deltas remain layout-sensitive.

Alloc/call remains 0. Wire/hash identities remain unchanged.

Run: https://github.com/SunSi12138/SharpLink/actions/runs/35684919723

Copy link
Copy Markdown
Owner Author

NativeAOT Serialize/Deserialize dispatch attribution follow-up on head ba849b9:

The earlier asymmetry (Serialize ~0%, Deserialize ~25-30%) does not survive controlled layout experiments.

Controls added:

  1. Deserialize return-used vs return-unused loops.
  2. Same-method dual-path loops where interface and concrete calls live in the same method.
  3. Two branch-order/layout variants (A/B).
  4. The exact same dispatch helper source is now overlaid into the dev baseline as well as the candidate before build.

Latest three-round same-machine results:

  • JIT + Dynamic PGO ON:

    • Serialize concrete vs interface: about -1%
    • Deserialize return-used: about -3%
    • Deserialize return-unused: ~0%
      => PGO erases most class-interface dispatch cost.
  • JIT + Dynamic PGO OFF:

    • Serialize: about -20% to -22%
    • Deserialize: about -20% to -21%
      => stable direct-call benefit, independent of return-value consumption.
  • NativeAOT isolated loops:

    • Serialize: roughly -44%
    • Deserialize return-used: roughly -40% to -45%
    • Deserialize return-unused: roughly -33% to -38%
  • NativeAOT same-method layout controls:

    • Serialize layout A: ~-25%
    • Serialize layout B: ~-45%
    • Deserialize layout A/B: ~-22% to -25%

Disassembly explains the structural difference:

  • interface loop: lea __InterfaceDispatchCell + call *(%r11) on every iteration;
  • concrete loop: one direct relative call to DispatchProbeCodec.Serialize/Deserialize.
    The interface loop is also larger and uses an indirect branch. The A/B variation, especially Serialize moving from ~25% to ~45% with only branch/layout ordering changed, shows the exact percentage is highly front-end/code-alignment sensitive under NativeAOT.

Return-used vs return-unused controls show the Serialize/Deserialize discrepancy was not caused by Deserialize's returned value or ABI. Both operations have the same underlying direct-call advantage. The reliable conclusion is qualitative/structural, not the exact 25/45% magnitude.

This microkernel remains an upper-bound isolation result because the probe method body is intentionally tiny and NoInlining. Real generated Codec bodies dilute the call cost; whole generated-codec/RPC evidence should still be used for production effect size.

Latest evidence run: https://github.com/SunSi12138/SharpLink/actions/runs/35684919723

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant