Repository navigation
perf: preserve concrete codec types in generated code - #727
Conversation
|
Performance follow-up after the custom-struct fix and same-binary controls:
Run: https://github.com/SunSi12138/SharpLink/actions/runs/35681317541 Interpretation: the concrete-dispatch codegen transition is real, but absolute dev-vs-candidate NativeAOT microbench deltas are sensitive to image layout. Same-binary controls are the safer evidence for dispatch cost. |
|
Added a #720-style isolated dispatch microkernel on head 4daa536. The timed sample calls one dedicated tight loop once; the loop itself performs 2,000,000 Codec calls in normal evidence runs. There is no per-item delegate/RPC/transport overhead. The probe is a sealed class implementing IRpcCodec; Serialize/Deserialize are NoInlining and do only minimal state work so the measurement isolates interface indirect call vs concrete sealed direct call. Three alternating same-machine rounds:
Alloc/call is 0 in all probe rows. Codegen confirms the intended shapes:
Interpretation matches #720's earlier signal: Dynamic PGO can erase most of the class-interface dispatch cost, while PGO-disabled JIT exposes a clear ~15-20% microkernel dispatch tax. NativeAOT also has a meaningful direct-call signal for this Deserialize shape. This is intentionally an upper/isolation measurement, not an end-to-end throughput claim; the generated DTO/RPC measurements remain the realistic-effect layer. Run: https://github.com/SunSi12138/SharpLink/actions/runs/35682751219 |
|
NativeAOT dispatch follow-up is now narrowed down with return-dependency and same-method layout controls (head ba849b9, run 35684919723). What changed in the probe:
Latest same-machine result:
Same-method layout controls:
This rules out the earlier hypotheses that the NativeAOT signal was primarily caused by ReadOnlySequence.Length, Deserialize return-value consumption, or simply having separate benchmark methods. The structural difference remains the call shape:
The right interpretation is therefore:
Alloc/call remains 0. Wire/hash identities remain unchanged. Run: https://github.com/SunSi12138/SharpLink/actions/runs/35684919723 |
|
NativeAOT Serialize/Deserialize dispatch attribution follow-up on head ba849b9: The earlier asymmetry (Serialize ~0%, Deserialize ~25-30%) does not survive controlled layout experiments. Controls added:
Latest three-round same-machine results:
Disassembly explains the structural difference:
Return-used vs return-unused controls show the Serialize/Deserialize discrepancy was not caused by Deserialize's returned value or ABI. Both operations have the same underlying direct-call advantage. The reliable conclusion is qualitative/structural, not the exact 25/45% magnitude. This microkernel remains an upper-bound isolation result because the probe method body is intentionally tiny and NoInlining. Real generated Codec bodies dilute the call cost; whole generated-codec/RPC evidence should still be used for production effect size. Latest evidence run: https://github.com/SunSi12138/SharpLink/actions/runs/35684919723 |
Implements #726.
This change preserves compile-time concrete Codec types in generated code when the final binding is statically fixed and legally nameable.
Key points:
Base: dev