Repository navigation
Conversation
The vectorized aggregate loops load eight vectors per iteration but fold them into one accumulator one at a time, so the loop runs at the latency of the aggregation operator rather than its throughput. Combine each group of four as a tree instead, for operators that declare reassociation exact, which today means integer addition and multiplication. Floating point keeps the original order, so its results are unchanged.
|
Azure Pipelines: Successfully started running 3 pipeline(s). 13 pipeline(s) were filtered out due to trigger conditions. There may be pipelines that require an authorized user to comment /azp run to run. |
Contributor
|
Tagging subscribers to this area: @dotnet/area-system-numerics |
Contributor
|
Tagging subscribers to this area: @dotnet/area-system-numerics-tensors |
Author
|
Built the runtime locally and ran the suite on this branch — |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
The vectorized aggregate loops in
Aggregateload eight vectors per iteration but fold them into a single accumulator one at a time:Every one of those depends on the previous, so the loop runs at the latency of the operator rather than its throughput, and the eightfold unrolling buys only the loads. On this machine
Sum<int>over a megabyte-sized span moves about 32 GB/s, far below what the loads alone would allow.How
Each group of four now goes through
AggregateFour, which combines the four vectors as a tree and folds the result into the accumulator once, leaving two dependent operations per iteration instead of eight. The tail (the masked first and last blocks and thecase 7 ... case 0jump table) is untouched, so spans shorter than eight vectors take exactly the same path as before.Reordering is only correct where the operator is associative over
T, so it is gated by a newstatic virtual bool CanReassociate => false;onIAggregationOperator<T>, overridden byAddOperator<T>andMultiplyOperator<T>withtypeof(T) != typeof(float) && typeof(T) != typeof(double). Those two are the only operators used in the aggregation position of this method. For integers, unchecked addition and multiplication are exact inZ/2^n, so results are bitwise identical, overflow included. When the flag is false,AggregateFourperforms the original four operations in the original order, sofloatanddoubleresults cannot move. The default is false so that an operator that is not associative inT-HalfAsInt16AggregationOperator, which hasT = shortbut computes infloat, is an example already in this assembly - does not opt in by accident.Floating point deliberately gets nothing here. Any reassociation changes the rounding, and
Sum's documentation only allows results to differ between platforms, not between versions on one platform. If you would like it behind an opt-in as well, say so and I will follow up.Test Plan
System.Numerics.Tensorstests pass: 5656 total, 0 failed (osx-arm64,dotnet build /t:Test -c Release -f net11.0).Added
Sum_MatchesScalarSumAcrossUnrolledBlocksandSumOfSquares_MatchesScalarSumAcrossUnrolledBlockstoGenericIntegerTensorPrimitivesTests, which compare against a scalar wrapping sum at lengths that straddle the eight-vector loop and its jump-table tail (1, 3, 8, 15, 16, 17, ... 1025, 4099).Differential against a build of unmodified
main: 4000 random cases per element type forint,uint,long,ulong,short,byte,floatanddouble, covering every API that reaches this code (Sum,SumOfSquares,SumOfMagnitudes,Dot,Product,ProductOfSums,ProductOfDifferences, and for floating pointDistance,Norm,CosineSimilarity). All results identical.A second, independently written harness compared the two builds over ten element types, lengths 1 to 600 plus 65536 and 1048576, four value regimes including
MinValue/MaxValue, NaN, infinities, negative zero and denormals, with floating-point results compared as raw bit patterns and thrownOverflowExceptions compared as results: about 230,000 results per build, zero differences.Benchmarks
osx-arm64 (
Vector128; this machine has no AVX2 or AVX-512), ns/op, alternating the two builds round by round and taking the minimum, since the box was not perfectly quiet:As bandwidth, every integer
Sum/SumOfSquarescase at 1M elements moves about 31-33 GB/s before and 108-112 GB/s after, across four element widths - the loop is memory bound now instead of dependency bound.floatanddoublesit at about 30 GB/s on both builds.The multiply-based aggregates gain less because the multiply, not the accumulator chain, is the limit:
Dot<int>1.77x,Dot<byte>1.14x,Dot<long>andSumOfSquares<long>flat. The floating-point rows above are noise around 1.00 - they take the unchanged path - and the two lowest (0.91, 0.96) did not reproduce in a second session.I do not have x64 hardware, so the
Vector256andVector512variants of these loops are covered by the tests and the differential but not by measurements. Inputs shorter than eight vectors are unchanged by construction.