Repository navigation
[release/11.0] Backport Tensor correctness and shape consistency fixes - #135281
Merged
tannergooding merged 2 commits intoOct 6, 2026
Merged
tannergooding merged 2 commits into
tannergooding merged 2 commits into
Conversation
…34914) ## Summary - Fix Tensor shape, storage, broadcasting, overlap, resize, and enumeration correctness; cover the behavior with regression tests and document intentional differences from NumPy. Empty shapes use zero automatically computed strides so unreachable products cannot overflow before the zero-length dimension. - Dispatch elementwise Tensor operations to TensorPrimitives for dense layouts and contiguous trailing slices, including compatible outer-dimension broadcasts. Preserve indexed traversal for incompatible layouts and ordering-sensitive reductions. - Accelerate strided flattening, equality, non-dense random fills, and non-dense Resize without changing logical iteration order. - Correct vector tangent range reduction discovered by CI: use guaranteed fused multiply-add on FMA hardware and an input-dependent scalar fallback for cancellation-sensitive vectors on other hardware, retaining vector evaluation for safe inputs. ## Bug reports and fuzzing findings - `ReadOnlyTensorSpan<T>` now rejects incompatible `System.Array` element types (dotnet#133459), and `Tensor<T>` slice pinning starts at the slice offset (dotnet#134691). - `ResizeTo` fills unwritten logical destination elements with `default(T)` and rejects ambiguous overlapping destination layouts; the behavior is documented (dotnet#128555). - Cropped `TensorSpan<T>.FlattenTo` now copies contiguous rows rather than individual elements, addressing the reported regression (dotnet#121463); the exact issue repro is measured below. - This PR addresses the `Tensor` findings [3, 6, 7, 8, 9, and 13 in the community SharpFuzz report](https://github.com/Mrnikbobjeff/runtime/blob/combined-sharpfuzz-net11-findings/src/libraries/Fuzzing/Findings/FINDINGS.md): strided `ToString`, broadcast pairwise comparisons, empty `*Any` comparisons, preserving the element when squeezing all singleton dimensions, `ResizeTo` zero-fill, and constructor/slice exception contracts. Squeezing retains a rank-one `[1]` shape rather than NumPy's rank-zero shape; this intentional difference is documented in the library README. Findings 4, 5, 10, 11, and 12 concern `TensorPrimitives` and are not addressed here. - **Partially addresses dotnet#125663:** contiguous non-dense slices avoid per-element copying in several operations, but the broader non-dense performance and allocation work remains open. ## Measurements Local BenchmarkDotNet on Windows x64 (Ryzen 9 7950X, .NET 11). Representative results below compare the prior implementation to this branch except where noted: | Scenario | Prior | Current | | --- | ---: | ---: | | Dense Add (64x16) | 12.69 us | 0.121 us | | Gapped-row Add (64x16) | 11.58 us | 2.29 us | | Broadcast-row Add (64x16) | 12.58 us | 1.91 us | | Transposed Add (64x16) | 12.54 us | 11.60 us | | Gapped-row FlattenTo | 5.66 us | 0.79 us | | Gapped-row SequenceEqual | 6.41 us | 0.57 us | | dotnet#121463 cropped FlattenTo repro (3x512x512 bytes) | 2.030 ms | 19.71 us | The dotnet#121463 row uses the issue's original 3x2000x2048 input cropped to 3x512x512, run locally against a saved pre-optimization assembly and this PR's Release assembly on the same .NET 11 host (approximately 103x faster). Additional local comparisons against equivalent indexed loops: non-dense Resize 1,909 -> 289 ns; uniform fill 5.56 -> 2.41 us; Gaussian fill 13.97 -> 10.88 us. Unary float Sin versus equivalent scalar iteration: dense 2.88 -> 0.63 us, gapped rows 8.41 -> 3.15 us. No additional managed allocation was measured in these samples. The transposed case is not addressed by the trailing-slice optimization. Following review, sliced elementwise dispatch is gated centrally at 32 elements, or 16 for tensor/tensor binary operations. Against the previously published PR head, the threshold follow-up improved four-element gapped Add from 237 to 135 ns and Abs from 134 to 68 ns; sixteen-element Add measured 235 to 243 ns, and retained sliced paths had small differences in both directions. Normal-path pool returns are retained without exceptional-path `finally` cleanup. Removing the reinstated `finally` blocks changed permutation codegen: rank-three permutation measured about 45 to 59 ns in the focused comparison, with unchanged allocations. These follow-up changes are not claimed as universally performance-neutral. The empty-stride correction was also compared locally with the preceding PR head across 12 dense-construction and reshape cases (ranks 1, 3, and 6, empty and nonempty). Allocation counts were unchanged; throughput measurements were noisy across repeated runs, so no performance improvement is claimed for this correctness fix. CI exposed tangent range-reduction errors in the newly activated vector path. Without fused multiply-add, `Tan(-32.986717f)` returned about -148998.98 instead of the scalar/reference result -177349.88. The correction is in `TensorPrimitives.Tan`, preserving Tensor's vector dispatch and existing tangent forwarding tolerances. FMA hardware uses guaranteed fused reduction rather than `MultiplyAddEstimate`, whose fused behavior is not guaranteed across runtimes. Other hardware retains non-fused vector evaluation when `abs(f) >= dn / 256`; otherwise the affected vector uses the existing scalar fallback. The threshold scales with the reduction count and has a documented rounding-error bound. With FMA unavailable, dense 4x1024 float Tan measured 4.31 -> 4.70 us on random inputs in [-1, 1] and 4.42 -> 8.08 us on [-50, 50], comparing the original vector implementation with the accuracy correction. Against a local global-FMA-gate implementation that scalarizes every non-FMA input, the corresponding results were 11.03 -> 4.70 us and 13.73 -> 8.08 us. Allocation counts remain unchanged. On FMA hardware, float and double kernels at all three widths retain the original instruction counts and code sizes; register allocation differs. FMA throughput measurements varied with code placement, so no precise performance-neutrality claim is made. ## Validation - System.Numerics.Tensors.Tests: 6,100 net11.0 and 200 net481 tests passed in each of five local Windows x64 configurations: default intrinsics, FMA-enabled 256-bit vectors, FMA-enabled 128-bit vectors, non-FMA 128-bit vectors, and intrinsics disabled. ARM64/Mono and Linux CI have not yet validated the tangent correction. - System.Numerics.Tensors Release analyzer build succeeded with zero warnings or errors. - Direct float/double vector-operator probes covered 1,966,080 checks per FMA configuration across all three widths, small and large random inputs, broad exponents, and homogeneous cancellation-boundary vectors. No tolerance failures occurred with or without FMA; non-FMA 256/512-bit operators were exercised through their software vector implementations on this host. - Independent deterministic layout oracle: 28,049 checks passed across 1,500 cases covering sliced, permuted, gapped, broadcast, and empty layouts. Resolves dotnet#128555 Resolves dotnet#133459 Resolves dotnet#134691 Resolves dotnet#121463 > [!NOTE] > This pull request description was generated by GitHub Copilot. --------- Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> (cherry picked from commit 137a670)
## Summary Make tensor shape handling consistent across broadcasting, elementwise operations, copies, equality, reshaping, axis operations, and views: - Treat default rank-zero empty values as effective shape `[0]`, without changing their stored `Rank`, `Lengths`, or `Strides`. Only redundant leading singleton axes can be added or removed; explicit zero axes and non-leading singleton axes remain significant. Stack and concatenate interpret axes using the first input's effective shape. - Fix all ten internal `Any` span kernels to traverse the input length and handle empty inputs correctly. - Keep empty view origins within source storage, handle zero-containing shape products without intermediate overflow, and retain native-sized counts and offsets instead of narrowing them prematurely. - Preserve overlap-safe ordering for native-width dense copies. Index-of-min/max reductions reuse `TensorPrimitives` for dense spans, dense suffixes, and bounded gathered blocks, preserving first-NaN, tie, signed-zero, and magnitude semantics. - Share concatenation validation and copying, avoid revalidating prepared stack inputs, and remove redundant shape comparisons after alignment. Caller-supplied destinations retain shape and overlap validation before writes. This includes behavioral compatibility changes, not new public API. Some formerly inconsistent shape combinations are now accepted or rejected according to the shared rules. For example, default empty values can explicitly broadcast to `[2, 0]`, while binary operations with `[0, 2]` reject the incompatible trailing dimension. Unsupported overlapping layouts and self-overlapping output layouts throw before writing. The README and affected API remarks document the shape and storage contracts; breaking-change documentation is required after merge. ## Validation - Release build of the tensor solution passed across its target frameworks. - Full suites passed normally and with hardware intrinsics disabled: **6,266 net11.0 tests and 200 net481 tests** per run, with no failures or skips. - Regression coverage includes default/ranked empties, leading-padding alignment, dense and strided destinations, rejection before writes, native-width logical counts, forced chunk boundaries, NaNs/ties, high ranks, and protected-memory boundaries. - Two consecutive independent shape/storage audit passes covered **149,246 cases each**, with no failures; the second used software intrinsics. ## Performance Local BenchmarkDotNet comparisons used Windows x64, a Ryzen 9 7950X, and the .NET 11 RC SDK/runtime. For a native-width broadcast view with shape `[nint.MaxValue / 1024, 1024]`, strides `[0, 1]`, and a NaN at the end of the first dense suffix, `IndexOfMax` measured **295 ns** with dense-suffix primitive dispatch versus **4,050 ns** with an elementwise native-width fallback. This compares dispatch strategies and early NaN termination, not full traversal of a multi-gigabyte allocation or throughput against upstream's length-narrowing path. Concatenation, stack, and dispatch comparisons retained the same allocation counts. Separate-process timings varied substantially and some apparent regressions reversed with run order. A same-process comparison using separately loaded assemblies measured row dispatch at **326 -> 309 ns** for rank 2 and **351 -> 356 ns** for rank 6, within measurement variability. No general throughput improvement is claimed for the validation simplification or metadata snapshots. > [!NOTE] > This pull request was authored with assistance from GitHub Copilot. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> (cherry picked from commit be7b579)
|
Azure Pipelines: Successfully started running 4 pipeline(s). 12 pipeline(s) were filtered out due to trigger conditions. There may be pipelines that require an authorized user to comment /azp run to run. |
Contributor
|
Tagging subscribers to this area: @dotnet/area-system-numerics-tensors |
artl93
approved these changes
Oct 6, 2026
artl93
left a comment
Member
There was a problem hiding this comment.
Tensors - relatively new. Math and bounds should work. Approved.
This was referenced Oct 6, 2026
This was referenced Oct 6, 2026
Member
Author
|
/ba-g unrelated CI failures |
Member
Author
|
CI failures are unrelated, is HTTP/3 cookie handling #133842, Mono/browser Double min/max NaN behavior https://github.com/dotnet/runtime/issues/133311], and Windows x86 Checked Process test hangs #135001 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes Issue dotnet/runtime#133459, dotnet/runtime#134691, dotnet/runtime#128555, and dotnet/runtime#121463.
main PR dotnet/runtime#134914 and dotnet/runtime#135060.
Description
Backport both merged Tensor fixes to
release/11.0in their original order:137a6701515dd7427e78bc46bcdd520c8d6e3a76: fix array element-type validation, slice pinning, shape/storage validation, broadcasting, overlap handling, resize zero-fill, enumeration, and numerical operations. Restore efficient dense and contiguous-slice dispatch, including the tangent range-reduction accuracy correction needed by that dispatch.be7b579e4230520b5c47cd956513147e13f737a6: make shape alignment consistent across operations, treat default rank-zero empties as effective shape[0]without changing their stored metadata, fix all ten internalAnyspan kernels, and preserve native-width counts, overlap-safe copy ordering, and valid empty-view origins.Both cherry-picks applied without conflicts or backport-specific source changes. The combined implementation, regression tests, README, and API remarks match the merged changes; release-specific compatibility suppressions are left unchanged. No public API is added.
Customer Impact
Without these fixes, incompatible
System.Arraystorage can be interpreted as the wrong element type and read beyond its bounds. Pinning a tensor slice can pass the parent's data, rather than the slice's data, to native consumers such as inference engines. Other affected operations can produce incorrect results, mishandle empty or broadcast shapes, omit the required resize zero-fill, or corrupt overlapping output.The existing cropped/gapped
FlattenToperformance regression and per-element overhead in dense and contiguous-slice operations also remain. Native-backed spans with more thanint.MaxValuelogical elements retain paths that prematurely narrow counts or offsets. Taking both fixes together keeps the initial correctness/dispatch changes paired with the subsequent shape-consistency and native-width corrections.Regression
Yes, in part, but these are not all newly introduced .NET 11 regressions. The array type-safety issue and cropped
FlattenToregression were introduced by the .NET 10 Tensor rewrite in dotnet/runtime#114927 and remain inrelease/11.0. Slice pinning was reproduced with the 10.0.9 package. The remaining changes address existing correctness and consistency defects rather than one recent servicing regression.Testing
Local Windows x64 validation on the actual
release/11.0backport:.\build.cmd clr+libs -rc release -lc releaseon the unmodified release baseline: passed, zero warnings/errors..\dotnet.cmd build src\libraries\System.Numerics.Tensors\System.Numerics.Tensors.slnx -c Release /p:RuntimeConfiguration=Releaseafter both cherry-picks: passed, zero warnings/errors. This built the configurednet11.0,net10.0,netstandard2.0, andnet462implementations and both test targets..\dotnet.cmd build src\libraries\System.Numerics.Tensors\tests\System.Numerics.Tensors.Tests.csproj -t:Test -c Release /p:RuntimeConfiguration=Release: passed in the default configuration./p:testnobuild=true --no-restorewas rerun under each environment override below: all passed. Fresh XML results were checked for the full counts, with no failures, errors, or skips.DOTNET_PreferredVectorBitWidth=256DOTNET_PreferredVectorBitWidth=128DOTNET_EnableAVX2=0DOTNET_EnableHWIntrinsic=0Coverage includes incompatible arrays, pinning, dense/strided/broadcast layouts, rejection before writes, empty shapes/views, high ranks, native-width logical counts, forced copy/reduction chunk boundaries, NaNs/ties/signed zero, and protected-memory boundaries.
Local BenchmarkDotNet ShortRun comparisons used the same built Release runtime with distinct baseline/backport Tensor assemblies on a Ryzen 9 7950X. Assembly loading and the active vector/FMA configuration were verified. The 10 default-intrinsic jobs and four non-FMA tangent jobs all produced measured results:
[32, 1][-1, 1], FMA[-50, 50], FMANo managed allocation was measured in these samples. Small timing differences are not claimed as precise gains or proof of performance neutrality; the non-FMA wide-range slowdown is an accuracy tradeoff, discussed below. These checks do not replace the more extensive measurements in the main PRs.
Linux/macOS, Arm64, Mono, NativeAOT, and browser/WASM were not locally validated.
Risk
Low-to-moderate risk. The changes are consistent backports of two fixes already merged on main, applied without conflicts or release-specific source adaptations. The extensive regression tests passed in all five local configurations, and the scope is confined to
System.Numerics.Tensors. Remaining risk is in the intentional behavior changes and edge cases described below, rather than the patch size.Intentional behavior changes include accepting/rejecting shape combinations consistently, interpreting default empties as effective shape
[0], rejecting incompatible array element types and unsupported overlapping layouts before writes, and zero-filling newly exposed resized elements. Consumers relying on the previous inconsistent behavior can be affected. There is no compatibility switch.On hardware without FMA, cancellation-sensitive tangent vectors now use scalar evaluation to avoid inaccurate results. The wide-range local sample increased from 5.60 to 9.19 us, approximately 64%; this is an acknowledged throughput cost for correctness, not a performance-neutral change. FMA-capable hardware retains fused vector range reduction.
Breaking-change documentation is already tracked by dotnet/docs#56290 and dotnet/docs#56305. Both currently name .NET 12 Preview 1; their shipment-version metadata needs to be aligned with the approved .NET 11 delivery if this backport is accepted.
Package authoring no longer needed in .NET 9
IMPORTANT: Starting with .NET 9, you no longer need to edit a NuGet package's csproj to enable building and bump the version.
Keep in mind that we still need package authoring in .NET 8 and older versions.
Note
This pull request description was generated with assistance from GitHub Copilot.