Skip to content

[release/11.0] Backport Tensor correctness and shape consistency fixes - #135281

Merged
tannergooding merged 2 commits into
dotnet:release/11.0from
tannergooding:tannergooding-combined-11-0-backport
Oct 6, 2026
Merged

tannergooding merged 2 commits into
dotnet:release/11.0from
tannergooding:tannergooding-combined-11-0-backport

Conversation

@tannergooding

@tannergooding tannergooding commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Fixes Issue dotnet/runtime#133459, dotnet/runtime#134691, dotnet/runtime#128555, and dotnet/runtime#121463.

main PR dotnet/runtime#134914 and dotnet/runtime#135060.

Description

Backport both merged Tensor fixes to release/11.0 in their original order:

  • 137a6701515dd7427e78bc46bcdd520c8d6e3a76: fix array element-type validation, slice pinning, shape/storage validation, broadcasting, overlap handling, resize zero-fill, enumeration, and numerical operations. Restore efficient dense and contiguous-slice dispatch, including the tangent range-reduction accuracy correction needed by that dispatch.
  • be7b579e4230520b5c47cd956513147e13f737a6: make shape alignment consistent across operations, treat default rank-zero empties as effective shape [0] without changing their stored metadata, fix all ten internal Any span kernels, and preserve native-width counts, overlap-safe copy ordering, and valid empty-view origins.

Both cherry-picks applied without conflicts or backport-specific source changes. The combined implementation, regression tests, README, and API remarks match the merged changes; release-specific compatibility suppressions are left unchanged. No public API is added.

Customer Impact

Without these fixes, incompatible System.Array storage can be interpreted as the wrong element type and read beyond its bounds. Pinning a tensor slice can pass the parent's data, rather than the slice's data, to native consumers such as inference engines. Other affected operations can produce incorrect results, mishandle empty or broadcast shapes, omit the required resize zero-fill, or corrupt overlapping output.

The existing cropped/gapped FlattenTo performance regression and per-element overhead in dense and contiguous-slice operations also remain. Native-backed spans with more than int.MaxValue logical elements retain paths that prematurely narrow counts or offsets. Taking both fixes together keeps the initial correctness/dispatch changes paired with the subsequent shape-consistency and native-width corrections.

Regression

Yes, in part, but these are not all newly introduced .NET 11 regressions. The array type-safety issue and cropped FlattenTo regression were introduced by the .NET 10 Tensor rewrite in dotnet/runtime#114927 and remain in release/11.0. Slice pinning was reproduced with the 10.0.9 package. The remaining changes address existing correctness and consistency defects rather than one recent servicing regression.

Testing

Local Windows x64 validation on the actual release/11.0 backport:

  • .\build.cmd clr+libs -rc release -lc release on the unmodified release baseline: passed, zero warnings/errors.
  • .\dotnet.cmd build src\libraries\System.Numerics.Tensors\System.Numerics.Tensors.slnx -c Release /p:RuntimeConfiguration=Release after both cherry-picks: passed, zero warnings/errors. This built the configured net11.0, net10.0, netstandard2.0, and net462 implementations and both test targets.
  • .\dotnet.cmd build src\libraries\System.Numerics.Tensors\tests\System.Numerics.Tensors.Tests.csproj -t:Test -c Release /p:RuntimeConfiguration=Release: passed in the default configuration.
  • The same full test command with /p:testnobuild=true --no-restore was rerun under each environment override below: all passed. Fresh XML results were checked for the full counts, with no failures, errors, or skips.
Configuration Environment override net11.0 passed net481 passed
Default intrinsics None 6,266 200
256-bit vectors DOTNET_PreferredVectorBitWidth=256 6,266 200
128-bit vectors DOTNET_PreferredVectorBitWidth=128 6,266 200
AVX2/FMA disabled DOTNET_EnableAVX2=0 6,266 200
All intrinsics disabled DOTNET_EnableHWIntrinsic=0 6,266 200

Coverage includes incompatible arrays, pinning, dense/strided/broadcast layouts, rejection before writes, empty shapes/views, high ranks, native-width logical counts, forced copy/reduction chunk boundaries, NaNs/ties/signed zero, and protected-memory boundaries.

Local BenchmarkDotNet ShortRun comparisons used the same built Release runtime with distinct baseline/backport Tensor assemblies on a Ryzen 9 7950X. Assembly loading and the active vector/FMA configuration were verified. The 10 default-intrinsic jobs and four non-FMA tangent jobs all produced measured results:

Scenario Release baseline Combined backport
Dense float Add, 64x16 16.37 us 0.114 us
Gapped-row float Add, 64x16, strides [32, 1] 16.25 us 2.05 us
Gapped-row float FlattenTo, same layout 7.92 us 0.875 us
Float Tan, 4,096 values in [-1, 1], FMA 1.80 us 1.72 us
Float Tan, 4,096 values in [-50, 50], FMA 1.63 us 1.53 us
Float Tan, same small range, no FMA 4.98 us 5.08 us
Float Tan, same wide range, no FMA 5.60 us 9.19 us

No managed allocation was measured in these samples. Small timing differences are not claimed as precise gains or proof of performance neutrality; the non-FMA wide-range slowdown is an accuracy tradeoff, discussed below. These checks do not replace the more extensive measurements in the main PRs.

Linux/macOS, Arm64, Mono, NativeAOT, and browser/WASM were not locally validated.

Risk

Low-to-moderate risk. The changes are consistent backports of two fixes already merged on main, applied without conflicts or release-specific source adaptations. The extensive regression tests passed in all five local configurations, and the scope is confined to System.Numerics.Tensors. Remaining risk is in the intentional behavior changes and edge cases described below, rather than the patch size.

Intentional behavior changes include accepting/rejecting shape combinations consistently, interpreting default empties as effective shape [0], rejecting incompatible array element types and unsupported overlapping layouts before writes, and zero-filling newly exposed resized elements. Consumers relying on the previous inconsistent behavior can be affected. There is no compatibility switch.

On hardware without FMA, cancellation-sensitive tangent vectors now use scalar evaluation to avoid inaccurate results. The wide-range local sample increased from 5.60 to 9.19 us, approximately 64%; this is an acknowledged throughput cost for correctness, not a performance-neutral change. FMA-capable hardware retains fused vector range reduction.

Breaking-change documentation is already tracked by dotnet/docs#56290 and dotnet/docs#56305. Both currently name .NET 12 Preview 1; their shipment-version metadata needs to be aligned with the approved .NET 11 delivery if this backport is accepted.

Package authoring no longer needed in .NET 9

IMPORTANT: Starting with .NET 9, you no longer need to edit a NuGet package's csproj to enable building and bump the version.
Keep in mind that we still need package authoring in .NET 8 and older versions.

Note

This pull request description was generated with assistance from GitHub Copilot.

tannergooding and others added 2 commits October 6, 2026 07:59
…34914)

## Summary
- Fix Tensor shape, storage, broadcasting, overlap, resize, and
enumeration correctness; cover the behavior with regression tests and
document intentional differences from NumPy. Empty shapes use zero
automatically computed strides so unreachable products cannot overflow
before the zero-length dimension.
- Dispatch elementwise Tensor operations to TensorPrimitives for dense
layouts and contiguous trailing slices, including compatible
outer-dimension broadcasts. Preserve indexed traversal for incompatible
layouts and ordering-sensitive reductions.
- Accelerate strided flattening, equality, non-dense random fills, and
non-dense Resize without changing logical iteration order.
- Correct vector tangent range reduction discovered by CI: use
guaranteed fused multiply-add on FMA hardware and an input-dependent
scalar fallback for cancellation-sensitive vectors on other hardware,
retaining vector evaluation for safe inputs.

## Bug reports and fuzzing findings
- `ReadOnlyTensorSpan<T>` now rejects incompatible `System.Array`
element types (dotnet#133459), and `Tensor<T>` slice pinning
starts at the slice offset (dotnet#134691).
- `ResizeTo` fills unwritten logical destination elements with
`default(T)` and rejects ambiguous overlapping destination layouts; the
behavior is documented (dotnet#128555).
- Cropped `TensorSpan<T>.FlattenTo` now copies contiguous rows rather
than individual elements, addressing the reported regression
(dotnet#121463); the exact issue repro is measured below.
- This PR addresses the `Tensor` findings [3, 6, 7, 8, 9, and 13 in the
community SharpFuzz
report](https://github.com/Mrnikbobjeff/runtime/blob/combined-sharpfuzz-net11-findings/src/libraries/Fuzzing/Findings/FINDINGS.md):
strided `ToString`, broadcast pairwise comparisons, empty `*Any`
comparisons, preserving the element when squeezing all singleton
dimensions, `ResizeTo` zero-fill, and constructor/slice exception
contracts. Squeezing retains a rank-one `[1]` shape rather than NumPy's
rank-zero shape; this intentional difference is documented in the
library README. Findings 4, 5, 10, 11, and 12 concern `TensorPrimitives`
and are not addressed here.
- **Partially addresses dotnet#125663:** contiguous non-dense
slices avoid per-element copying in several operations, but the broader
non-dense performance and allocation work remains open.

## Measurements
Local BenchmarkDotNet on Windows x64 (Ryzen 9 7950X, .NET 11).
Representative results below compare the prior implementation to this
branch except where noted:

| Scenario | Prior | Current |
| --- | ---: | ---: |
| Dense Add (64x16) | 12.69 us | 0.121 us |
| Gapped-row Add (64x16) | 11.58 us | 2.29 us |
| Broadcast-row Add (64x16) | 12.58 us | 1.91 us |
| Transposed Add (64x16) | 12.54 us | 11.60 us |
| Gapped-row FlattenTo | 5.66 us | 0.79 us |
| Gapped-row SequenceEqual | 6.41 us | 0.57 us |
| dotnet#121463 cropped FlattenTo repro (3x512x512 bytes) |
2.030 ms | 19.71 us |

The dotnet#121463 row uses the issue's original 3x2000x2048
input cropped to 3x512x512, run locally against a saved pre-optimization
assembly and this PR's Release assembly on the same .NET 11 host
(approximately 103x faster). Additional local comparisons against
equivalent indexed loops: non-dense Resize 1,909 -> 289 ns; uniform fill
5.56 -> 2.41 us; Gaussian fill 13.97 -> 10.88 us. Unary float Sin versus
equivalent scalar iteration: dense 2.88 -> 0.63 us, gapped rows 8.41 ->
3.15 us. No additional managed allocation was measured in these samples.
The transposed case is not addressed by the trailing-slice optimization.

Following review, sliced elementwise dispatch is gated centrally at 32
elements, or 16 for tensor/tensor binary operations. Against the
previously published PR head, the threshold follow-up improved
four-element gapped Add from 237 to 135 ns and Abs from 134 to 68 ns;
sixteen-element Add measured 235 to 243 ns, and retained sliced paths
had small differences in both directions. Normal-path pool returns are
retained without exceptional-path `finally` cleanup. Removing the
reinstated `finally` blocks changed permutation codegen: rank-three
permutation measured about 45 to 59 ns in the focused comparison, with
unchanged allocations. These follow-up changes are not claimed as
universally performance-neutral.

The empty-stride correction was also compared locally with the preceding
PR head across 12 dense-construction and reshape cases (ranks 1, 3, and
6, empty and nonempty). Allocation counts were unchanged; throughput
measurements were noisy across repeated runs, so no performance
improvement is claimed for this correctness fix.

CI exposed tangent range-reduction errors in the newly activated vector
path. Without fused multiply-add, `Tan(-32.986717f)` returned about
-148998.98 instead of the scalar/reference result -177349.88. The
correction is in `TensorPrimitives.Tan`, preserving Tensor's vector
dispatch and existing tangent forwarding tolerances. FMA hardware uses
guaranteed fused reduction rather than `MultiplyAddEstimate`, whose
fused behavior is not guaranteed across runtimes. Other hardware retains
non-fused vector evaluation when `abs(f) >= dn / 256`; otherwise the
affected vector uses the existing scalar fallback. The threshold scales
with the reduction count and has a documented rounding-error bound.

With FMA unavailable, dense 4x1024 float Tan measured 4.31 -> 4.70 us on
random inputs in [-1, 1] and 4.42 -> 8.08 us on [-50, 50], comparing the
original vector implementation with the accuracy correction. Against a
local global-FMA-gate implementation that scalarizes every non-FMA
input, the corresponding results were 11.03 -> 4.70 us and 13.73 -> 8.08
us. Allocation counts remain unchanged. On FMA hardware, float and
double kernels at all three widths retain the original instruction
counts and code sizes; register allocation differs. FMA throughput
measurements varied with code placement, so no precise
performance-neutrality claim is made.

## Validation
- System.Numerics.Tensors.Tests: 6,100 net11.0 and 200 net481 tests
passed in each of five local Windows x64 configurations: default
intrinsics, FMA-enabled 256-bit vectors, FMA-enabled 128-bit vectors,
non-FMA 128-bit vectors, and intrinsics disabled. ARM64/Mono and Linux
CI have not yet validated the tangent correction.
- System.Numerics.Tensors Release analyzer build succeeded with zero
warnings or errors.
- Direct float/double vector-operator probes covered 1,966,080 checks
per FMA configuration across all three widths, small and large random
inputs, broad exponents, and homogeneous cancellation-boundary vectors.
No tolerance failures occurred with or without FMA; non-FMA 256/512-bit
operators were exercised through their software vector implementations
on this host.
- Independent deterministic layout oracle: 28,049 checks passed across
1,500 cases covering sliced, permuted, gapped, broadcast, and empty
layouts.

Resolves dotnet#128555
Resolves dotnet#133459
Resolves dotnet#134691
Resolves dotnet#121463

> [!NOTE]
> This pull request description was generated by GitHub Copilot.

---------

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
(cherry picked from commit 137a670)
## Summary

Make tensor shape handling consistent across broadcasting, elementwise
operations, copies, equality, reshaping, axis operations, and views:

- Treat default rank-zero empty values as effective shape `[0]`, without
changing their stored `Rank`, `Lengths`, or `Strides`. Only redundant
leading singleton axes can be added or removed; explicit zero axes and
non-leading singleton axes remain significant. Stack and concatenate
interpret axes using the first input's effective shape.
- Fix all ten internal `Any` span kernels to traverse the input length
and handle empty inputs correctly.
- Keep empty view origins within source storage, handle zero-containing
shape products without intermediate overflow, and retain native-sized
counts and offsets instead of narrowing them prematurely.
- Preserve overlap-safe ordering for native-width dense copies.
Index-of-min/max reductions reuse `TensorPrimitives` for dense spans,
dense suffixes, and bounded gathered blocks, preserving first-NaN, tie,
signed-zero, and magnitude semantics.
- Share concatenation validation and copying, avoid revalidating
prepared stack inputs, and remove redundant shape comparisons after
alignment. Caller-supplied destinations retain shape and overlap
validation before writes.

This includes behavioral compatibility changes, not new public API. Some
formerly inconsistent shape combinations are now accepted or rejected
according to the shared rules. For example, default empty values can
explicitly broadcast to `[2, 0]`, while binary operations with `[0, 2]`
reject the incompatible trailing dimension. Unsupported overlapping
layouts and self-overlapping output layouts throw before writing. The
README and affected API remarks document the shape and storage
contracts; breaking-change documentation is required after merge.

## Validation

- Release build of the tensor solution passed across its target
frameworks.
- Full suites passed normally and with hardware intrinsics disabled:
**6,266 net11.0 tests and 200 net481 tests** per run, with no failures
or skips.
- Regression coverage includes default/ranked empties, leading-padding
alignment, dense and strided destinations, rejection before writes,
native-width logical counts, forced chunk boundaries, NaNs/ties, high
ranks, and protected-memory boundaries.
- Two consecutive independent shape/storage audit passes covered
**149,246 cases each**, with no failures; the second used software
intrinsics.

## Performance

Local BenchmarkDotNet comparisons used Windows x64, a Ryzen 9 7950X, and
the .NET 11 RC SDK/runtime.

For a native-width broadcast view with shape `[nint.MaxValue / 1024,
1024]`, strides `[0, 1]`, and a NaN at the end of the first dense
suffix, `IndexOfMax` measured **295 ns** with dense-suffix primitive
dispatch versus **4,050 ns** with an elementwise native-width fallback.
This compares dispatch strategies and early NaN termination, not full
traversal of a multi-gigabyte allocation or throughput against
upstream's length-narrowing path.

Concatenation, stack, and dispatch comparisons retained the same
allocation counts. Separate-process timings varied substantially and
some apparent regressions reversed with run order. A same-process
comparison using separately loaded assemblies measured row dispatch at
**326 -> 309 ns** for rank 2 and **351 -> 356 ns** for rank 6, within
measurement variability. No general throughput improvement is claimed
for the validation simplification or metadata snapshots.

> [!NOTE]
> This pull request was authored with assistance from GitHub Copilot.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
(cherry picked from commit be7b579)
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 4 pipeline(s).
12 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-numerics-tensors
See info in area-owners.md if you want to be subscribed.

@artl93 artl93 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tensors - relatively new. Math and bounds should work. Approved.

@tannergooding

Copy link
Copy Markdown
Member Author

/ba-g unrelated CI failures

@tannergooding

Copy link
Copy Markdown
Member Author

CI failures are unrelated, is HTTP/3 cookie handling #133842, Mono/browser Double min/max NaN behavior https://github.com/dotnet/runtime/issues/133311], and Windows x86 Checked Process test hangs #135001

@tannergooding
tannergooding merged commit 1ee1930 into dotnet:release/11.0 Oct 6, 2026
113 of 121 checks passed
@tannergooding
tannergooding deleted the tannergooding-combined-11-0-backport branch October 6, 2026 21:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants