Skip to content

Combine aggregate vectors as a tree for associative operators - #134327

Closed
tahakocal wants to merge 1 commit into
dotnet:mainfrom
tahakocal:perf/tensors-aggregate-tree
Closed

tahakocal wants to merge 1 commit into
dotnet:mainfrom
tahakocal:perf/tensors-aggregate-tree

Conversation

@tahakocal

Copy link
Copy Markdown

Why

The vectorized aggregate loops in Aggregate load eight vectors per iteration but fold them into a single accumulator one at a time:

vresult = TAggregationOperator.Invoke(vresult, vector1);
vresult = TAggregationOperator.Invoke(vresult, vector2);
vresult = TAggregationOperator.Invoke(vresult, vector3);
vresult = TAggregationOperator.Invoke(vresult, vector4);

Every one of those depends on the previous, so the loop runs at the latency of the operator rather than its throughput, and the eightfold unrolling buys only the loads. On this machine Sum<int> over a megabyte-sized span moves about 32 GB/s, far below what the loads alone would allow.

How

Each group of four now goes through AggregateFour, which combines the four vectors as a tree and folds the result into the accumulator once, leaving two dependent operations per iteration instead of eight. The tail (the masked first and last blocks and the case 7 ... case 0 jump table) is untouched, so spans shorter than eight vectors take exactly the same path as before.

Reordering is only correct where the operator is associative over T, so it is gated by a new static virtual bool CanReassociate => false; on IAggregationOperator<T>, overridden by AddOperator<T> and MultiplyOperator<T> with typeof(T) != typeof(float) && typeof(T) != typeof(double). Those two are the only operators used in the aggregation position of this method. For integers, unchecked addition and multiplication are exact in Z/2^n, so results are bitwise identical, overflow included. When the flag is false, AggregateFour performs the original four operations in the original order, so float and double results cannot move. The default is false so that an operator that is not associative in T - HalfAsInt16AggregationOperator, which has T = short but computes in float, is an example already in this assembly - does not opt in by accident.

Floating point deliberately gets nothing here. Any reassociation changes the rounding, and Sum's documentation only allows results to differ between platforms, not between versions on one platform. If you would like it behind an opt-in as well, say so and I will follow up.

Test Plan

System.Numerics.Tensors tests pass: 5656 total, 0 failed (osx-arm64, dotnet build /t:Test -c Release -f net11.0).

Added Sum_MatchesScalarSumAcrossUnrolledBlocks and SumOfSquares_MatchesScalarSumAcrossUnrolledBlocks to GenericIntegerTensorPrimitivesTests, which compare against a scalar wrapping sum at lengths that straddle the eight-vector loop and its jump-table tail (1, 3, 8, 15, 16, 17, ... 1025, 4099).

Differential against a build of unmodified main: 4000 random cases per element type for int, uint, long, ulong, short, byte, float and double, covering every API that reaches this code (Sum, SumOfSquares, SumOfMagnitudes, Dot, Product, ProductOfSums, ProductOfDifferences, and for floating point Distance, Norm, CosineSimilarity). All results identical.

A second, independently written harness compared the two builds over ten element types, lengths 1 to 600 plus 65536 and 1048576, four value regimes including MinValue/MaxValue, NaN, infinities, negative zero and denormals, with floating-point results compared as raw bit patterns and thrown OverflowExceptions compared as results: about 230,000 results per build, zero differences.

Benchmarks

osx-arm64 (Vector128; this machine has no AVX2 or AVX-512), ns/op, alternating the two builds round by round and taking the minimum, since the box was not perfectly quiet:

API T N before after ratio
Sum int 65536 8092 2361 3.43x
Sum int 1048576 130839 37387 3.50x
Sum long 1048576 261453 75950 3.44x
Sum short 1048576 65008 18605 3.49x
Sum byte 1048576 33268 9454 3.52x
SumOfSquares int 1048576 131041 37915 3.46x
SumOfSquares byte 1048576 33386 9403 3.55x
Dot int 1048576 131682 74230 1.77x
Dot short 65536 4039 2393 1.69x
Dot long 1048576 288897 294067 0.98x
Sum float 1048576 142126 147433 0.96x
Sum double 1048576 277493 304239 0.91x
SumOfSquares double 1048576 309870 308743 1.00x
Dot float 1048576 161097 155100 1.04x

As bandwidth, every integer Sum/SumOfSquares case at 1M elements moves about 31-33 GB/s before and 108-112 GB/s after, across four element widths - the loop is memory bound now instead of dependency bound. float and double sit at about 30 GB/s on both builds.

The multiply-based aggregates gain less because the multiply, not the accumulator chain, is the limit: Dot<int> 1.77x, Dot<byte> 1.14x, Dot<long> and SumOfSquares<long> flat. The floating-point rows above are noise around 1.00 - they take the unchanged path - and the two lowest (0.91, 0.96) did not reproduce in a second session.

I do not have x64 hardware, so the Vector256 and Vector512 variants of these loops are covered by the tests and the differential but not by measurements. Inputs shorter than eight vectors are unchanged by construction.

The vectorized aggregate loops load eight vectors per iteration but fold
them into one accumulator one at a time, so the loop runs at the latency
of the aggregation operator rather than its throughput. Combine each
group of four as a tree instead, for operators that declare reassociation
exact, which today means integer addition and multiplication. Floating
point keeps the original order, so its results are unchanged.
@dotnet-policy-service dotnet-policy-service Bot added the community-contribution Indicates that the PR has been added by a community member label Sep 21, 2026
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 3 pipeline(s).
13 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-numerics
See info in area-owners.md if you want to be subscribed.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @dotnet/area-system-numerics-tensors
See info in area-owners.md if you want to be subscribed.

@tahakocal

Copy link
Copy Markdown
Author

Built the runtime locally and ran the suite on this branch — System.Numerics.Tensors.Tests, osx-arm64 Release: 5,656 total, 0 errors, 0 failed, 0 skipped.

@tahakocal tahakocal closed this Sep 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area-System.Numerics.Tensors community-contribution Indicates that the PR has been added by a community member

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants