Repository navigation
Conversation
|
Weird - the new tests don't fail on my M5 Max. Do they fail on your machine if you remove the patch in |
|
I mean the |
|
ok so I got the first case "row stride >= 2^16" always fail on master, are you having the same result? the second case (conv2d case) is non-deterministic because on master it reads OOB, uninitialized/random memory |
|
On diff --git a/tests/test-backend-ops.cpp b/tests/test-backend-ops.cpp
index 08c29eec6..101126531 100644
--- a/tests/test-backend-ops.cpp
+++ b/tests/test-backend-ops.cpp
@@ -9082,6 +9082,14 @@ static std::vector<std::unique_ptr<test_case>> make_test_cases_eval() {
}
}
+ // permuted src1 with a row stride >= 2^16 elements, as produced by the MLA + FA attention
+ // epilogue with 128 heads (e.g. deepseek32, dots3note)
+ test_cases.emplace_back(new test_mul_mat(GGML_TYPE_F16, GGML_TYPE_F32, 128, 32, 512, {128, 1}, {1, 1}, {0, 2, 1, 3}));
+
+ // K not a multiple of the mat-mat tile size, as produced by conv_2d im2col with K = 14*14*3
+ // (vision patch embedding, see #25652)
+ test_cases.emplace_back(new test_mul_mat(GGML_TYPE_F16, GGML_TYPE_F32, 64, 32, 588, {1, 1}, {1, 1}));
+
// BF16 is absent from base_types: add the 3 standard non-contig permutations explicitly
test_cases.emplace_back(new test_mul_mat(GGML_TYPE_BF16, GGML_TYPE_F32, 16, 1, 256, {2, 3}, {1, 1}, {0, 2, 1, 3}));
test_cases.emplace_back(new test_mul_mat(GGML_TYPE_BF16, GGML_TYPE_F32, 16, 1, 256, {2, 3}, {1, 1}, {0, 1, 3, 2}));
@@ -9711,12 +9719,13 @@ static std::vector<std::unique_ptr<test_case>> make_test_cases_eval() {
if (nh == 1 && hsk != 320 && hsk != 576) continue;
for (int nr3 : { 1, 3, }) {
if (hsk > 64 && nr3 > 1) continue; // skip broadcast for large head sizes
- for (int nr2 : { 1, 4, 8, 12, 16, 20, 32 }) {
+ for (int nr2 : { 1, 4, 8, 12, 16, 20, 32, 128 }) {
if (nr2 == 8 && hsk != 192) continue;
if (nr2 == 12 && hsk != 128) continue;
if (nr2 == 16 && hsk != 192) continue;
if (nr2 == 20 && (nh != 1 || hsk != 576)) continue;
if (nr2 == 32 && (nh != 1 || hsk != 320)) continue;
+ if (nr2 == 128 && (nh != 1 || hsk != 576)) continue; // deepseek32/dots3note MLA-as-MQA (128 q heads, 1 kv head)
//for (int kv : { 1, 17, 31, 33, 61, 113, 65, 127, 129, 130, 255, 260, 371, 380, 407, 512, 1024, }) {
for (int kv : { 113, 512, 1024, }) {
if (nr2 != 1 && kv != 512) continue;make -j && ./bin/test-backend-ops -b MTL0 -o MUL_MAT
ggml_metal_library_init: using embedded metal library
ggml_metal_library_init: loaded in 6.287 sec
ggml_metal_rsets_init: creating a residency set collection (keep_alive = 180 s)
ggml_metal_device_init: GPU name: MTL0 (Apple M5 Max)
ggml_metal_device_init: GPU family: MTLGPUFamilyApple10 (1010)
ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003)
ggml_metal_device_init: GPU family: MTLGPUFamilyMetal4 (5002)
ggml_metal_device_init: simdgroup reduction = true
ggml_metal_device_init: simdgroup matrix mul. = true
ggml_metal_device_init: has unified memory = true
ggml_metal_device_init: has bfloat = true
ggml_metal_device_init: has tensor = true
ggml_metal_device_init: use residency sets = true
ggml_metal_device_init: use shared buffers = true
ggml_metal_device_init: recommendedMaxWorkingSetSize = 40200.90 MB
Testing 3 devices
ggml_metal_init: allocating
ggml_metal_init: found device: Apple M5 Max
ggml_metal_init: picking default device: Apple M5 Max
ggml_metal_init: use fusion = true
ggml_metal_init: use concurrency = true
ggml_metal_init: use graph optimize = true
Backend 1/3: MTL0
Device description: Apple M5 Max
Device memory: 38338 MB (38337 MB free)
...
1156/1156 tests passed
Backend MTL0: OK |
|
I got this on my machine: Running on macOS 26.3.2 / 25D2150 |
|
Here is the full system info section I think this bug is quite tricky as it's non-deterministic. Lmk any other info you need to debug this: |
| // - the last K tile is read out of bounds when ne00 is not a multiple of the tile size, | ||
| // unlike src0 which is staged through threadgroup memory with zero padding | ||
| // example: the im2col src1 of a conv_2d with K = 14*14*3 (see #25652) | ||
| (!props_dev->has_tensor || (nb11/ggml_type_size(op->src[1]->type) < 65536 && ne00 % 32 == 0))) { |
There was a problem hiding this comment.
@ngxson Could you try only adding the ne00 % 32 condition without the nb11 limit (I can't find any information that the stride is limited to 16 bits):
| (!props_dev->has_tensor || (nb11/ggml_type_size(op->src[1]->type) < 65536 && ne00 % 32 == 0))) { | |
| (!props_dev->has_tensor || (ne00 % 32 == 0))) { |
Let me know if this passes all the tests (+ mtmd) on your machine.
There was a problem hiding this comment.
Ok so the test backend op case fails, but mtmd test is OK

Overview
Fix 2 cases spotted on:
Fix #25652
Tested locally:
@ggerganov some comments can be redundant, feel free to edit or remove them
Requirements