Skip to content

vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 - #27952

Merged
0cc4m merged 3 commits into
masterfrom
0cc4m/vulkan-coopmat-int8
Sep 24, 2026
Merged

0cc4m merged 3 commits into
masterfrom
0cc4m/vulkan-coopmat-int8

Conversation

@0cc4m

@0cc4m 0cc4m commented Aug 29, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Currently with tensor/matrix cores we only use fp16 instructions, so dequant to shared memory -> fp16 GEMM. But many devices have higher throughput int8 tensor/matrix cores, so similar to the existing DP4A-based scalar MMQ path it would be good to have a coopmat int8 MMQ path. I have tried to make this work in the past, but failed due to architectural constraints in the coopmat extension.

But based on @pwilkin's ideas in #27493, I finally managed to get it to work with good performance on RDNA3 and 4. This new MMQ cm1 shader supports q4_0, q4_1, q5_0, q5_1, q8_0, q3_k, q4_k, q5_k, q6_k, mxfp4, nvfp4 and iq4_nl on RDNA3 and RDNA4. Strix Halo prompt processing performance is significantly improved, RDNA4 is more neutral, but MoE prompt processing is also up significantly. The shader is limited to those two architectures since it hardcodes their specific coopmat access patterns. More architectures can be added if needed.

q4_1, q5_1, q4_k, q5_k and nvfp4 on RDNA4 are disabled for MUL_MAT since they run slower than existing fp16 matmul. Additionally nvfp4 is also disabled for MUL_MAT_ID.

Benchmarks

AMD Radeon 8060S (Strix Halo, RDNA3.5)

MUL_MAT test-backend-ops perf geomean us (lower = better)

quant vk-master vk-feat rocm-master feat vs master feat vs rocm
q4_0 5183 4016 5064 1.29x 1.26x
q4_1 5123 4157 5189 1.23x 1.25x
q5_0 5881 4340 5273 1.35x 1.21x
q5_1 5574 4271 5362 1.31x 1.26x
q8_0 5645 4443 5421 1.27x 1.22x
iq4_nl 5471 4535 5061 1.21x 1.12x
mxfp4 5205 4329 4903 1.20x 1.13x
q3_K 6352 5718 5974 1.11x 1.04x
q4_K 5535 4749 5304 1.17x 1.12x
q5_K 5637 4940 5437 1.14x 1.10x
q6_K 6545 5747 5323 1.14x 0.93x
nvfp4 5262 5430 6290 0.97x 1.16x

MUL_MAT_ID test-backend-ops perf geomean us (lower = better)

quant vk-master vk-feat rocm-master feat vs master feat vs rocm
q4_0 867 658 817 1.32x 1.24x
q4_1 869 663 875 1.31x 1.32x
q5_0 990 707 799 1.40x 1.13x
q5_1 962 691 785 1.39x 1.14x
q8_0 843 692 853 1.22x 1.23x
iq4_nl 951 723 747 1.32x 1.03x
mxfp4 925 703 727 1.32x 1.03x
q3_K 1138 856 853 1.33x 1.00x
q4_K 990 779 762 1.27x 0.98x
q5_K 994 824 844 1.21x 1.02x
q6_K 1094 899 1089 1.22x 1.21x
nvfp4 969 822 1121 1.18x 1.36x

llama-bench pp512 t/s (higher = better)

model quant kind vk-master vk-feat rocm-master feat vs master feat vs rocm
Meta-Llama-3-8B Q4_0 dense 862.7 1075.7 924.2 1.25x 1.16x
Meta-Llama-3-8B Q4_1 dense 937.3 1073.2 897.9 1.15x 1.20x
Meta-Llama-3.1-8B IQ4_NL dense 901.7 1035.1 934.3 1.15x 1.11x
Meta-Llama-3.1-8B Q3_K_S dense 791.3 867.3 830.0 1.10x 1.04x
Meta-Llama-3.1-8B Q4_K_S dense 874.3 975.1 907.1 1.12x 1.08x
Meta-Llama-3.1-8B Q6_K dense 772.2 862.9 729.4 1.12x 1.18x
Meta-Llama-3.1-8B Q8_0 dense 823.4 1030.5 874.4 1.25x 1.18x
llama2-13b-tiefighter Q5_K_M dense 515.0 547.1 519.6 1.06x 1.05x
Qwen3.8-27B Q4_K_M dense 186.5 254.1 237.6 1.36x 1.07x
gpt-oss-20b MXFP4 MoE 1071.6 1410.5 1005.8 1.32x 1.40x
Qwen3.6-35B-A3B Q4_0 MoE 883.7 1175.9 743.4 1.33x 1.58x
Qwen3.6-35B-A3B Q4_K_M MoE 777.8 1023.6 693.9 1.32x 1.48x
gemma-4-26B-A4B Q6_K MoE 737.2 925.5 723.1 1.26x 1.28x
chart_pp512
AMD Radeon AI PRO R9700 (RDNA4)

MUL_MAT test-backend-ops perf geomean us (lower = better)

quant vk-master vk-feat rocm-master feat vs master feat vs rocm RDNA4 MMQ
q4_0 1146 1120 1036 1.02x 0.93x on
q4_1 1123 1123 1189 1.00x 1.06x off (fallback)
q5_0 1378 1234 1116 1.12x 0.90x on
q5_1 1272 1273 1203 1.00x 0.94x off (fallback)
q8_0 1324 1179 1024 1.12x 0.87x on
iq4_nl 1304 1198 958 1.09x 0.80x on
mxfp4 1326 1202 960 1.10x 0.80x on
q3_K 1974 1616 1466 1.22x 0.91x on
q4_K 1372 1373 1153 1.00x 0.84x off (fallback)
q5_K 1433 1435 1180 1.00x 0.82x off (fallback)
q6_K 1746 1583 2395 1.10x 1.51x on
nvfp4 1330 1332 1554 1.00x 1.17x off (fallback)

MUL_MAT_ID test-backend-ops perf geomean us (lower = better)

quant vk-master vk-feat rocm-master feat vs master feat vs rocm RDNA4 MMQ
q4_0 210 243 202 0.86x 0.83x on
q4_1 205 253 219 0.81x 0.87x on
q5_0 249 265 210 0.94x 0.79x on
q5_1 227 260 219 0.87x 0.84x on
q8_0 228 252 194 0.90x 0.77x on
iq4_nl 237 259 195 0.92x 0.75x on
mxfp4 241 261 195 0.93x 0.75x on
q3_K 364 329 247 1.11x 0.75x on
q4_K 251 307 216 0.82x 0.71x on
q5_K 261 316 218 0.83x 0.69x on
q6_K 312 327 343 0.95x 1.05x on
nvfp4 242 243 282 1.00x 1.16x off (fallback)

llama-bench pp512 t/s (higher = better)

model quant kind vk-master vk-feat rocm-master feat vs master feat vs rocm
Llama-3-8B Q4_0 dense 4656 4584 4370 0.98x 1.05x
Llama-3-8B Q4_1 dense 4751 4771 3911 1.00x 1.22x
Llama-3-8B Q4_K_S dense 3881 3968 3887 1.02x 1.02x
Llama-3.1-8B IQ4_NL dense 4059 4169 4488 1.03x 0.93x
Llama-3.1-8B Q3_K_S dense 2882 3228 3321 1.12x 0.97x
Llama-3.1-8B Q6_K dense 3208 3259 2165 1.02x 1.51x
Llama-3.1-8B Q8_0 dense 4230 4378 4473 1.03x 0.98x
Mistral-Nemo-12B Q4_0 dense 2925 2882 2876 0.99x 1.00x
gemma-4-12B Q4_0 dense 2761 2711 2429 0.98x 1.12x
Qwen3.8-27B Q6_K dense 955 1040 764 1.09x 1.36x
gpt-oss-20B MXFP4 MoE 3936 5631 4231 1.43x 1.33x
Qwen3.6-35B-A3B Q4_0 MoE 3196 4046 2638 1.27x 1.53x
Qwen3.6-35B-A3B Q4_K_M MoE 2760 3365 2240 1.22x 1.50x
Muse-Glimmer-30B Q6_K MoE 995 1065 792 1.07x 1.34x
chart_pp512

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, AI was used for optimization loops. Code and performance was manually checked/verified.

@github-actions github-actions Bot added Vulkan Issues specific to the Vulkan backend ggml changes relating to the ggml tensor library for machine learning labels Aug 29, 2026
@0cc4m
0cc4m marked this pull request as ready for review August 29, 2026 12:10
@0cc4m
0cc4m requested a review from a team as a code owner August 29, 2026 12:10
#define ACC_BIAS_F 12582912.0f
const bool USE_MAGIC_BIAS = WARP != 32;

// Accumulator row for element e: RDNA4 blocked, RDNA3/3.5 interleaved.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are these layouts entirely based on the hardware definition, or do they depend partly on the compiler. For example, at some point I changed the NVIDIA compiler so the layout of a 16x16 matrix in four registers was permuted to 0,2,1,3 vs earlier compiler versions.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think it's a little bit of both. I have confirmed that Linux' RADV driver based this layout on what the hardware requires, so it should be fixed. I have also checked that at least RDNA3.5 uses the same layout on the Windows driver, and this branch does improve performance there too.

I think we can replace this with a generic solution once enough drivers support coopmat maintenance1.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On the other PR, you mentioned that you couldn't probe the layout because it used too many registers, but if the matrix is 16x16 then we only need 4 bits each for row+col and that's only a couple registers for wave32. I'm surprised storing those would cause a perf issue.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I removed it in a number of optimizations that reduced overall register use. It's possible I can readd it now, but I don't think it's worth the complexity with coopmat maintenance1 on the way.

return;
}
#else
// L2-friendly workgroup scheduling

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would be interesting to separate out what perf gain is from this vs using int8. I've been tempted to do something like this for the other shaders, but with ubatch=512 still being the default it usually doesn't matter much (probably helps more for stable diffusion kind of workloads)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very small, hard to measure. Maybe 2% on RDNA3.5, but my Strix Halo laptop has a lot of thermal noise in benchmarks, so not sure. Might not be worth it. I didn't measure a difference on RDNA4.

@edt-xx

edt-xx commented Sep 2, 2026

Copy link
Copy Markdown

Built b10760 + this pr. Using a RX7900XT (RDNA3) with RADV. Running usloth/Qwen3.8-27B-Q4_0.gguf before applying this PR I was seeing:

[35963] 0.10.092.723 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =   4096, progress = 0.05, t =   4.25 s / 963.84 tokens per second
...
[35963] 2.42.829.150 I slot print_timing: id  0 | task 0 | prompt processing, n_tokens =  75995, progress = 1.00, t = 156.13 s / 486.73 tokens per second

after the same sort of prefill gets (+8% at start to +3% at end of prefill)

[40189] 181.52.418.736 I slot print_timing: id  0 | task 2568 | prompt processing, n_tokens =   4096, progress = 0.05, t =   3.92 s / 1046.10 tokens per second
...
[40189] 184.24.107.728 I slot print_timing: id  0 | task 2568 | prompt processing, n_tokens =  77865, progress = 1.00, t = 156.04 s / 498.99 tokens per second

There is too much variance in token gen to say anything. This has positive effects in a real use case.

@NickM-27

NickM-27 commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

Apologies if this is not helpful / spam, but I have been running this on my 7900XTX for a few days and it has been a big improvement, so wanted to share a before / after llama bench run:

master:

ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = AMD Radeon RX 7900 XTX (RADV NAVI31) (radv) | uma: 0 | fp16: 1 | bf16: 0 | fp4: 0 | warp size: 64 | shared memory: 65536 | int dot: 1 | matrix cores: KHR_coopmat
load_backend: loaded Vulkan backend from /app/libggml-vulkan.so
load_backend: loaded CPU backend from /app/libggml-cpu-alderlake.so
| model                          |       size |     params | backend    | ngl | n_ubatch |  fa |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | --: | --------------: | -------------------: |
| gemma4 26B.A4B Q4_0            |  13.26 GiB |    25.23 B | Vulkan     |  -1 |     1024 |   1 |           pp512 |      3410.53 ± 22.72 |
| gemma4 26B.A4B Q4_0            |  13.26 GiB |    25.23 B | Vulkan     |  -1 |     1024 |   1 |           tg128 |        135.64 ± 0.80 |
| gemma4 26B.A4B Q4_0            |  13.26 GiB |    25.23 B | Vulkan     |  -1 |     1024 |   1 |   pp512 @ d8192 |      2477.27 ± 76.78 |
| gemma4 26B.A4B Q4_0            |  13.26 GiB |    25.23 B | Vulkan     |  -1 |     1024 |   1 |   tg128 @ d8192 |        120.42 ± 0.46 |
| gemma4 26B.A4B Q4_0            |  13.26 GiB |    25.23 B | Vulkan     |  -1 |     1024 |   1 |  pp512 @ d16384 |      2130.36 ± 35.16 |
| gemma4 26B.A4B Q4_0            |  13.26 GiB |    25.23 B | Vulkan     |  -1 |     1024 |   1 |  tg128 @ d16384 |        118.42 ± 0.14 |

this PR:

ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = AMD Radeon RX 7900 XTX (RADV NAVI31) (radv) | uma: 0 | fp16: dot2 | bf16: 1 | fp4: 0 | warp size: 64 | shared memory: 65536 | int dot: 1 | matrix cores: KHR_coopmat
load_backend: loaded Vulkan backend from /app/libggml-vulkan.so
load_backend: loaded CPU backend from /app/libggml-cpu-alderlake.so
| model                          |       size |     params | backend    | ngl | n_ubatch |  fa |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | --: | --------------: | -------------------: |
| gemma4 26B.A4B Q4_0            |  13.26 GiB |    25.23 B | Vulkan     |  -1 |     1024 |   1 |           pp512 |     4331.21 ± 132.67 |
| gemma4 26B.A4B Q4_0            |  13.26 GiB |    25.23 B | Vulkan     |  -1 |     1024 |   1 |           tg128 |        142.76 ± 0.92 |
| gemma4 26B.A4B Q4_0            |  13.26 GiB |    25.23 B | Vulkan     |  -1 |     1024 |   1 |   pp512 @ d8192 |      2885.67 ± 80.24 |
| gemma4 26B.A4B Q4_0            |  13.26 GiB |    25.23 B | Vulkan     |  -1 |     1024 |   1 |   tg128 @ d8192 |        125.29 ± 0.14 |
| gemma4 26B.A4B Q4_0            |  13.26 GiB |    25.23 B | Vulkan     |  -1 |     1024 |   1 |  pp512 @ d16384 |      2402.13 ± 44.16 |
| gemma4 26B.A4B Q4_0            |  13.26 GiB |    25.23 B | Vulkan     |  -1 |     1024 |   1 |  tg128 @ d16384 |        120.82 ± 0.18 |

@0cc4m

0cc4m commented Sep 7, 2026

Copy link
Copy Markdown
Contributor Author

@jeffbolznv Can you take a look at this? I know you can't test it, but at least to check the code is not breaking anything else.

#define ACC_BIAS_F 12582912.0f
const bool USE_MAGIC_BIAS = WARP != 32;

// Accumulator row for element e: RDNA4 blocked, RDNA3/3.5 interleaved.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

On the other PR, you mentioned that you couldn't probe the layout because it used too many registers, but if the matrix is 16x16 then we only need 4 bits each for row+col and that's only a couple registers for wave32. I'm surprised storing those would cause a perf issue.


const uint CM_ELEMS = (TM * TN) / WARP;
#define ACC_BIAS_BITS 0x4B400000
#define ACC_BIAS_F 12582912.0f

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume this is one of these "add an integer to the mantissa of this float value" tricks, but I couldn't easily tell what's going on.

#include "mul_mmq_cm1_funcs.glsl"

void main() {
#if defined(DATA_A_IQ4_NL)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I assume this will get the spec constant treatment before it goes in?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I checked and this is not easily possible, because unlike the fp16 shaders which dequant into shared memory, the MMQ shaders have to have custom shmem layouts for the quantized data in shared memory. This would get us into trouble with shmem size tricks again that previously especially the Nvidia driver did not like. Do you see a way to handle this?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, I don't have any clever ideas for this.

Comment thread ggml/src/ggml-vulkan/ggml-vulkan.cpp Outdated
const uint32_t tk_m = device->coopmat_support ? device->coopmat_k : 1;
const uint32_t tk_s = device->coopmat_support ? device->coopmat_k : 1;

const uint32_t itm_l = device->coopmat_int_support ? device->coopmat_int_m : 4;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

codex had one interesting finding here:


  - [P1] Keep cooperative-matrix dimensions out of scalar MMQ tiles.
    ggml/src/ggml-vulkan/ggml-vulkan.cpp:4389 now uses coopmat_int_* for l/m/s_warptile_mmq_int. If integer cooperative matrices are advertised but the new AMD CM1 path is not
    selected, these dimensions reach the scalar mul_mmq.comp shader, where TM/TN/TK have different meanings. Preserve the old scalar values and use the queried dimensions only in
    *_mmq_cm1_int.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, that was not correct. Fixed.

@vlforit

vlforit commented Sep 9, 2026

Copy link
Copy Markdown

Please rebase on top of master head, there are merge conflicts

@0cc4m

0cc4m commented Sep 9, 2026

Copy link
Copy Markdown
Contributor Author

Please rebase on top of master head, there are merge conflicts

No, this is a PR, it exists to facilitate merging this into the master branch, not for you to merge and use. There's a plan to get this merged, I'm following it.

@vlforit

vlforit commented Sep 9, 2026

Copy link
Copy Markdown

Could you share the plan that you are following?

@0cc4m

0cc4m commented Sep 9, 2026

Copy link
Copy Markdown
Contributor Author

It touches a lot of code, and there will be more merge conflicts with other PRs that get merged first. Rebasing now is just a waste of effort.

@vlforit

vlforit commented Sep 9, 2026

Copy link
Copy Markdown

It touches a lot of code, and there will be more merge conflicts with other PRs that get merged first. Rebasing now is just a waste of effort.

It took me 25 seconds to do rebase to master head, it is less than your answer.

No, this is a PR, it exists to facilitate merging this into the master branch, not for you to merge and use. There's a plan to get this merged, I'm following it.

You didn't answer to simple question, where is that plan you are following ?

@0cc4m

0cc4m commented Sep 9, 2026

Copy link
Copy Markdown
Contributor Author

It's in my head. Let me do my job. Stop this discussion or you will get blocked.

NickM-27 added a commit to NickM-27/llama.cpp that referenced this pull request Sep 9, 2026
@pwilkin

pwilkin commented Sep 9, 2026

Copy link
Copy Markdown
Member

@vlforit did we miss you somehow becoming the main coordinator of llama.cpp? If not, please be respectful to maintainers instead of demanding they engage you on things that are purely your personal whims.

@vlforit

vlforit commented Sep 9, 2026

Copy link
Copy Markdown

It's in my head.

Oh, that explains a lot.

@ggml-org ggml-org temporarily blocked vlforit Sep 9, 2026
@0cc4m
0cc4m force-pushed the 0cc4m/vulkan-coopmat-int8 branch from 224807f to 8c7611a Compare September 10, 2026 06:35
@MatheusMenezes08

Copy link
Copy Markdown

Discrete RDNA3 data point, not a bisect. RX 7800 XT (Windows, AMD proprietary Vulkan, KHR_coopmat), Qwen3.6-35B-A3B Q4_K_M, same flags on c77ae69 vs 84e76d8 (this PR is in the second build, plus 145 other commits).pp512 1174 to 1151 t/s (flat, not the Strix Halo +32%). g32 without MTP 42.4 to 36.1 t/s. MTP n=3 tg128: 55.8 to 25.8 t/s at 48k, 39.2 to 27.1 t/s at 130k, acceptance still 1.00. Full writeup: #29410.

@Railway9784

Copy link
Copy Markdown

Post-merge data point from an RDNA3 APU (unified memory, so please read this as iGPU-specific — it matches the edt-xx RX 7900XT result above in direction but is a very different memory regime).

Hardware / software

  • Ryzen 5 7640HS, Radeon 760M (RADV PHOENIX, gfx1103, RDNA3), UMA carve-out, Mesa 26.0.8-1ubuntu0.3, kernel 7.0.0, Ubuntu 26.04
  • llama-bench device line: uma: 1 | fp16: 1 | bf16: 0 | fp4: 0 | warp size: 64 | shared memory: 65536 | int dot: 1 | matrix cores: KHR_coopmat
  • llama.cpp built from master 84e76d8a2 with -DGGML_VULKAN=ON -DGGML_NATIVE=ON -DCMAKE_BUILD_TYPE=Release, shared libs. Baseline build is 434ddbbc0.

Method — three binaries, identical build flags, benched interleaved in one session (all three in the same process conditions, so drift cancels):

  1. 434ddbbc0 — neither this PR nor vulkan: add IQ4_XS MMQ/MMV matmul kernels #28415
  2. 84e76d8a2 with 70c4e1582 reverted (git revert on the master tree; verified mul_mmq_cm1 gone and matmul_iq4_xs_q8_1 still present in the compiled tree) — vulkan: add IQ4_XS MMQ/MMV matmul kernels #28415 only
  3. 84e76d8a2 — both

The workloads are the two models this box actually serves: Gemma 4 E2B IQ4_XS (-p 3461 -n 1157, the real request shape) and Qwen3-Embedding-0.6B Q8_0 (prefill only). Flags: -ngl 99 -t 4 --flash-attn on -ctk q8_0 -ctv q8_0 -b 2048 -ub 1024 (embedder -t 2 -b 2048).

test 1) neither PR 2) #28415 only 3) #28415 + #27952
pp3461 IQ4_XS 613.8 629.6 779.8 (+27%)
pp32768 IQ4_XS 373.6 * 378.8 421.7 (+13%)
pp512 Q8_0 2302 2333 2734 (+19%)
pp2048 Q8_0 1957 1946 2226 (+14%)
tg1157 decode 41.7 41.5 41.5 (flat)

* the pp32768 baseline is from an earlier session; the other rows are same-session. Pass-to-pass spread on pp3461 is ±0.3%.

Result — on this APU essentially the entire prefill win comes from this PR's coopmat1 path. Reverting it gives back ~94% of the pp3461 gain; the generic int8-dot IQ4_XS MMQ/MMVQ kernels from #28415 measure ~+2.6% (pp3461), ~+1.4% (32K depth) and ~0% on the Q8_0 embedder. Decode is unchanged in every build, consistent with an iGPU being bandwidth-bound on generation.

Two notes that may be useful:

  • The coopmat1 path clearly engages and pays on this APU even though vulkaninfo enumerates no cooperative-matrix property configs on RADV/PHOENIX — so "not enumerable" here does not mean "not used".
  • Because the win is RDNA3/RDNA4 matrix-core dependent, it looks driver-sensitive; I can re-run after a Mesa bump if that is helpful.

Measured with llama-bench through a scripted, counter-balanced harness; raw per-build output retained.

}
#undef X_CM1

if (device->coopmat_int_support && (rdna3 || rdna4)) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Add definitions.
#if defined(GGML_VULKAN_INTEGER_DOT_GLSLC_SUPPORT)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, it doesn't use integer dot.

Comment thread ggml/src/ggml-vulkan/ggml-vulkan.cpp
@Ankk98

Ankk98 commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

I was also trying to make this work for qwen3.8 27B on strix halo but I was not able to make it work and while trying that found this PR. Glad to see it working.

Few learnings from the mistakes I made:

  • Did not see open PRs beforehand
  • Using wrong correctness test for Q5_K, expecting o/p to match stock byte for byte but there is a precision change. Should have used NMSE tolerance in test-backend-ops plus KL divergence .
  • Temporature based throttling is a big issue on my strix halo
  • Bug in IQ4_XS port

@Kezii

Kezii commented Sep 25, 2026

Copy link
Copy Markdown

I'm unfortunately seeing sub 1% improvements on RDNA4, do we need to enable something or I am just unlucky with Q5 weights?

@0cc4m

0cc4m commented Sep 25, 2026

Copy link
Copy Markdown
Contributor Author

See benchmarks in the PR description, RDNA4 improvements are mostly for MoE.

pratiknarola-t added a commit to tetherto/qvac-fabric-llm.cpp that referenced this pull request Oct 1, 2026
Ports upstream 70c4e15 (vulkan: int8 coopmat1 matmul implementation
for AMD RDNA3 and RDNA4, ggml-org#27952): the mul_mmq_cm1 shaders, AMD_RDNA4
detection, cm1 int8 shared-memory checks and warptiles, and the q8_1
path selection on coopmat_int_support.

Adapted to this tree:
- pipelines are created with the existing CREATE_MM-era helpers since
  the upstream pipeline-map refactor (91f6a6c) is not here
- the cm1 MUL_MAT_ID shader uses the non-hoisted row ids and push
  constants of this tree (hoisting, 90c26fc, is not here)
- WARP_SIZE_IDX taken from upstream 1945e09

Upstream e85e15c (ggml-org#29409) needs no change here: its guards cover the
Intel FA decode shaders this tree does not have, and the cm1 shaders
are only generated when glslc supports coopmat.
pratiknarola-t added a commit to tetherto/qvac-fabric-llm.cpp that referenced this pull request Oct 1, 2026
Ports upstream 70c4e15 (vulkan: int8 coopmat1 matmul implementation
for AMD RDNA3 and RDNA4, ggml-org#27952): the mul_mmq_cm1 shaders, AMD_RDNA4
detection, cm1 int8 shared-memory checks and warptiles, and the q8_1
path selection on coopmat_int_support.

Adapted to this tree:
- pipelines are created with the existing CREATE_MM-era helpers since
  the upstream pipeline-map refactor (91f6a6c) is not here
- the cm1 MUL_MAT_ID shader uses the non-hoisted row ids and push
  constants of this tree (hoisting, 90c26fc, is not here)
- WARP_SIZE_IDX taken from upstream 1945e09

Upstream e85e15c (ggml-org#29409) needs no change here: its guards cover the
Intel FA decode shaders this tree does not have, and the cm1 shaders
are only generated when glslc supports coopmat.
pratiknarola-t added a commit to tetherto/qvac-fabric-llm.cpp that referenced this pull request Oct 1, 2026
Ports upstream 70c4e15 (vulkan: int8 coopmat1 matmul implementation
for AMD RDNA3 and RDNA4, ggml-org#27952): the mul_mmq_cm1 shaders, AMD_RDNA4
detection, cm1 int8 shared-memory checks and warptiles, and the q8_1
path selection on coopmat_int_support.

Adapted to this tree:
- pipelines are created with the existing CREATE_MM-era helpers since
  the upstream pipeline-map refactor (91f6a6c) is not here
- the cm1 MUL_MAT_ID shader uses the non-hoisted row ids and push
  constants of this tree (hoisting, 90c26fc, is not here)
- WARP_SIZE_IDX taken from upstream 1945e09

Upstream e85e15c (ggml-org#29409) needs no change here: its guards cover the
Intel FA decode shaders this tree does not have, and the cm1 shaders
are only generated when glslc supports coopmat.
pratiknarola-t added a commit to tetherto/qvac-fabric-llm.cpp that referenced this pull request Oct 2, 2026
Ports upstream 70c4e15 (vulkan: int8 coopmat1 matmul implementation
for AMD RDNA3 and RDNA4, ggml-org#27952): the mul_mmq_cm1 shaders, AMD_RDNA4
detection, cm1 int8 shared-memory checks and warptiles, and the q8_1
path selection on coopmat_int_support.

The shaders are byte-identical to upstream, including the hoisted
MUL_MAT_ID row ids and push constants (90c26fc is in this tree). The
pipelines are created the way upstream does it, through
create_mm_pipelines with the tc_mmq_cm1_int and tc_mmq_cm1_int_k tile
tables.

Adapted to this tree:
- quantize_y also requires that src0 needs no reformat copy and that
  src1 is dim01-contiguous. This tree sends non-dim01-contiguous
  quantized src0 through ggml_vk_mul_mat_q_f16 for its quant->f16 copy,
  and ggml_is_contiguous accepts permuted size-1 dims, so without the
  old !y_non_contig term the q8_1 path met the reformat copies (wrong
  MUL_MAT results, and a dispatch of a never-requested copy pipeline).
  The src0 check also fixes int-dot devices without coopmat, where the
  parent already paired the q8_1 path with the src0 copy.
  test_mul_mat_permuted_src0 covers the src0 side.
- AMD_RDNA4 goes into the architecture enum in ggml-vulkan.cpp, since
  ggml-vulkan-types.h (f172be7) is not here
- the IQ4_NL and NVFP4 cases upstream adds to
  ggml_vk_matmul_int_shmem_support are left out: this tree has no
  int-dot q8_1 pipelines for those types, so they would be dead
- the cm1 q8_1 shaders get the device_prefix that the other matmul
  shaders in vulkan-shaders-gen use
- the cm1 q8_1 shaders are generated in the fp16 coopmat pass only.
  This tree also runs the matmul shader pass for f32 coopmat inputs
  (e0bee17), which would otherwise build fp32 q8_1 variants that no
  pipeline creates

Upstream e85e15c (ggml-org#29409) needs no change here: its guards cover the
Intel FA decode shaders this tree does not have, and the cm1 shaders
are only generated when glslc supports coopmat.
pratiknarola-t added a commit to tetherto/qvac-fabric-llm.cpp that referenced this pull request Oct 2, 2026
Ports upstream 70c4e15 (vulkan: int8 coopmat1 matmul implementation
for AMD RDNA3 and RDNA4, ggml-org#27952): the mul_mmq_cm1 shaders, AMD_RDNA4
detection, cm1 int8 shared-memory checks and warptiles, and the q8_1
path selection on coopmat_int_support.

The shaders are byte-identical to upstream, including the hoisted
MUL_MAT_ID row ids and push constants (90c26fc is in this tree). The
pipelines are created the way upstream does it, through
create_mm_pipelines with the tc_mmq_cm1_int and tc_mmq_cm1_int_k tile
tables.

Adapted to this tree:
- quantize_y also requires that src0 needs no reformat copy and that
  src1 is dim01-contiguous. This tree sends non-dim01-contiguous
  quantized src0 through ggml_vk_mul_mat_q_f16 for its quant->f16 copy,
  and ggml_is_contiguous accepts permuted size-1 dims, so without the
  old !y_non_contig term the q8_1 path met the reformat copies (wrong
  MUL_MAT results, and a dispatch of a never-requested copy pipeline).
  The src0 check also fixes int-dot devices without coopmat, where the
  parent already paired the q8_1 path with the src0 copy.
  test_mul_mat_permuted_src0 covers the src0 side.
- AMD_RDNA4 goes into the architecture enum in ggml-vulkan.cpp, since
  ggml-vulkan-types.h (f172be7) is not here
- the IQ4_NL and NVFP4 cases upstream adds to
  ggml_vk_matmul_int_shmem_support are left out: this tree has no
  int-dot q8_1 pipelines for those types, so they would be dead
- the cm1 q8_1 shaders get the device_prefix that the other matmul
  shaders in vulkan-shaders-gen use
- the cm1 q8_1 shaders are generated in the fp16 coopmat pass only.
  This tree also runs the matmul shader pass for f32 coopmat inputs
  (e0bee17), which would otherwise build fp32 q8_1 variants that no
  pipeline creates

Upstream e85e15c (ggml-org#29409) needs no change here: its guards cover the
Intel FA decode shaders this tree does not have, and the cm1 shaders
are only generated when glslc supports coopmat.
gianni-cor pushed a commit to tetherto/qvac-fabric-llm.cpp that referenced this pull request Oct 3, 2026
Ports upstream 70c4e15 (vulkan: int8 coopmat1 matmul implementation
for AMD RDNA3 and RDNA4, ggml-org#27952): the mul_mmq_cm1 shaders, AMD_RDNA4
detection, cm1 int8 shared-memory checks and warptiles, and the q8_1
path selection on coopmat_int_support.

The shaders are byte-identical to upstream, including the hoisted
MUL_MAT_ID row ids and push constants (90c26fc is in this tree). The
pipelines are created the way upstream does it, through
create_mm_pipelines with the tc_mmq_cm1_int and tc_mmq_cm1_int_k tile
tables.

Adapted to this tree:
- quantize_y also requires that src0 needs no reformat copy and that
  src1 is dim01-contiguous. This tree sends non-dim01-contiguous
  quantized src0 through ggml_vk_mul_mat_q_f16 for its quant->f16 copy,
  and ggml_is_contiguous accepts permuted size-1 dims, so without the
  old !y_non_contig term the q8_1 path met the reformat copies (wrong
  MUL_MAT results, and a dispatch of a never-requested copy pipeline).
  The src0 check also fixes int-dot devices without coopmat, where the
  parent already paired the q8_1 path with the src0 copy.
  test_mul_mat_permuted_src0 covers the src0 side.
- AMD_RDNA4 goes into the architecture enum in ggml-vulkan.cpp, since
  ggml-vulkan-types.h (f172be7) is not here
- the IQ4_NL and NVFP4 cases upstream adds to
  ggml_vk_matmul_int_shmem_support are left out: this tree has no
  int-dot q8_1 pipelines for those types, so they would be dead
- the cm1 q8_1 shaders get the device_prefix that the other matmul
  shaders in vulkan-shaders-gen use
- the cm1 q8_1 shaders are generated in the fp16 coopmat pass only.
  This tree also runs the matmul shader pass for f32 coopmat inputs
  (e0bee17), which would otherwise build fp32 q8_1 variants that no
  pipeline creates

Upstream e85e15c (ggml-org#29409) needs no change here: its guards cover the
Intel FA decode shaders this tree does not have, and the cm1 shaders
are only generated when glslc supports coopmat.
gianni-cor pushed a commit to tetherto/qvac-fabric-llm.cpp that referenced this pull request Oct 3, 2026
Ports upstream 70c4e15 (vulkan: int8 coopmat1 matmul implementation
for AMD RDNA3 and RDNA4, ggml-org#27952): the mul_mmq_cm1 shaders, AMD_RDNA4
detection, cm1 int8 shared-memory checks and warptiles, and the q8_1
path selection on coopmat_int_support.

The shaders are byte-identical to upstream, including the hoisted
MUL_MAT_ID row ids and push constants (90c26fc is in this tree). The
pipelines are created the way upstream does it, through
create_mm_pipelines with the tc_mmq_cm1_int and tc_mmq_cm1_int_k tile
tables.

Adapted to this tree:
- quantize_y also requires that src0 needs no reformat copy and that
  src1 is dim01-contiguous. This tree sends non-dim01-contiguous
  quantized src0 through ggml_vk_mul_mat_q_f16 for its quant->f16 copy,
  and ggml_is_contiguous accepts permuted size-1 dims, so without the
  old !y_non_contig term the q8_1 path met the reformat copies (wrong
  MUL_MAT results, and a dispatch of a never-requested copy pipeline).
  The src0 check also fixes int-dot devices without coopmat, where the
  parent already paired the q8_1 path with the src0 copy.
  test_mul_mat_permuted_src0 covers the src0 side.
- AMD_RDNA4 goes into the architecture enum in ggml-vulkan.cpp, since
  ggml-vulkan-types.h (f172be7) is not here
- the IQ4_NL and NVFP4 cases upstream adds to
  ggml_vk_matmul_int_shmem_support are left out: this tree has no
  int-dot q8_1 pipelines for those types, so they would be dead
- the cm1 q8_1 shaders get the device_prefix that the other matmul
  shaders in vulkan-shaders-gen use
- the cm1 q8_1 shaders are generated in the fp16 coopmat pass only.
  This tree also runs the matmul shader pass for f32 coopmat inputs
  (e0bee17), which would otherwise build fp32 q8_1 variants that no
  pipeline creates

Upstream e85e15c (ggml-org#29409) needs no change here: its guards cover the
Intel FA decode shaders this tree does not have, and the cm1 shaders
are only generated when glslc supports coopmat.
@Arvamer

Arvamer commented Oct 3, 2026

Copy link
Copy Markdown

This regressed RDNA4 for batch sizes above 4 on Gemma 3 31B QAT by ~40%. Before:

PP TG B N_KV T_PP s S_PP t/s T_TG s S_TG t/s T s S t/s
512 128 1 640 0.447 1144.86 4.260 30.04 4.708 135.95
512 128 2 1280 0.901 1137.12 4.487 57.05 5.388 237.58
512 128 4 2560 1.800 1137.85 6.108 83.83 7.908 323.73
512 128 6 3840 2.715 1131.42 5.430 141.44 8.145 471.45
512 128 8 5120 3.630 1128.34 6.561 156.08 10.191 502.40

After:

PP TG B N_KV T_PP s S_PP t/s T_TG s S_TG t/s T s S t/s
512 128 1 640 0.421 1217.30 4.247 30.14 4.668 137.11
512 128 2 1280 0.847 1209.32 4.478 57.17 5.325 240.39
512 128 4 2560 1.693 1209.94 6.101 83.92 7.794 328.48
512 128 6 3840 2.545 1207.00 8.653 88.75 11.199 342.90
512 128 8 5120 3.401 1204.33 11.325 90.42 14.726 347.69

(The command was llama-batched-bench -m /mnt/data1/models/llm/gemma-4-31B-it-qat-UD-Q4_K_XL.gguf -ngl 99 -fa on -c 8192 -npp 512 -ntg 128 -npl 1,2,4,6,8)

According to Opus, this is because since this MR RDNA4 is treated as separate arch, so some optimizations are no longer enabled. Specifically, this line:

    // RDNA3: above four columns, static 4 rows for all types bench faster than the default
    const bool is_rdna3 = device->vendor_id == VK_VENDOR_ID_AMD && device->architecture == AMD_RDNA3;

From ggml/src/ggml-vulkan/ggml-vulkan.cpp.

Changing it to also allow RDNA4 restores the performance.

@0cc4m

0cc4m commented Oct 4, 2026

Copy link
Copy Markdown
Contributor Author

You're right, I overlooked that. Will fix.

pratiknarola-t added a commit to tetherto/qvac-fabric-llm.cpp that referenced this pull request Oct 5, 2026
Ports upstream 70c4e15 (vulkan: int8 coopmat1 matmul implementation
for AMD RDNA3 and RDNA4, ggml-org#27952): the mul_mmq_cm1 shaders, AMD_RDNA4
detection, cm1 int8 shared-memory checks and warptiles, and the q8_1
path selection on coopmat_int_support.

The shaders are byte-identical to upstream, including the hoisted
MUL_MAT_ID row ids and push constants (90c26fc is in this tree). The
pipelines are created the way upstream does it, through
create_mm_pipelines with the tc_mmq_cm1_int and tc_mmq_cm1_int_k tile
tables.

Adapted to this tree:
- quantize_y also requires that src0 needs no reformat copy and that
  src1 is dim01-contiguous. This tree sends non-dim01-contiguous
  quantized src0 through ggml_vk_mul_mat_q_f16 for its quant->f16 copy,
  and ggml_is_contiguous accepts permuted size-1 dims, so without the
  old !y_non_contig term the q8_1 path met the reformat copies (wrong
  MUL_MAT results, and a dispatch of a never-requested copy pipeline).
  The src0 check also fixes int-dot devices without coopmat, where the
  parent already paired the q8_1 path with the src0 copy.
  test_mul_mat_permuted_src0 covers the src0 side.
- AMD_RDNA4 goes into the architecture enum in ggml-vulkan.cpp, since
  ggml-vulkan-types.h (f172be7) is not here
- the IQ4_NL and NVFP4 cases upstream adds to
  ggml_vk_matmul_int_shmem_support are left out: this tree has no
  int-dot q8_1 pipelines for those types, so they would be dead
- the cm1 q8_1 shaders get the device_prefix that the other matmul
  shaders in vulkan-shaders-gen use
- the cm1 q8_1 shaders are generated in the fp16 coopmat pass only.
  This tree also runs the matmul shader pass for f32 coopmat inputs
  (e0bee17), which would otherwise build fp32 q8_1 variants that no
  pipeline creates

Upstream e85e15c (ggml-org#29409) needs no change here: its guards cover the
Intel FA decode shaders this tree does not have, and the cm1 shaders
are only generated when glslc supports coopmat.
gianni-cor pushed a commit to tetherto/qvac-fabric-llm.cpp that referenced this pull request Oct 5, 2026
Ports upstream 70c4e15 (vulkan: int8 coopmat1 matmul implementation
for AMD RDNA3 and RDNA4, ggml-org#27952): the mul_mmq_cm1 shaders, AMD_RDNA4
detection, cm1 int8 shared-memory checks and warptiles, and the q8_1
path selection on coopmat_int_support.

The shaders are byte-identical to upstream, including the hoisted
MUL_MAT_ID row ids and push constants (90c26fc is in this tree). The
pipelines are created the way upstream does it, through
create_mm_pipelines with the tc_mmq_cm1_int and tc_mmq_cm1_int_k tile
tables.

Adapted to this tree:
- quantize_y also requires that src0 needs no reformat copy and that
  src1 is dim01-contiguous. This tree sends non-dim01-contiguous
  quantized src0 through ggml_vk_mul_mat_q_f16 for its quant->f16 copy,
  and ggml_is_contiguous accepts permuted size-1 dims, so without the
  old !y_non_contig term the q8_1 path met the reformat copies (wrong
  MUL_MAT results, and a dispatch of a never-requested copy pipeline).
  The src0 check also fixes int-dot devices without coopmat, where the
  parent already paired the q8_1 path with the src0 copy.
  test_mul_mat_permuted_src0 covers the src0 side.
- AMD_RDNA4 goes into the architecture enum in ggml-vulkan.cpp, since
  ggml-vulkan-types.h (f172be7) is not here
- the IQ4_NL and NVFP4 cases upstream adds to
  ggml_vk_matmul_int_shmem_support are left out: this tree has no
  int-dot q8_1 pipelines for those types, so they would be dead
- the cm1 q8_1 shaders get the device_prefix that the other matmul
  shaders in vulkan-shaders-gen use
- the cm1 q8_1 shaders are generated in the fp16 coopmat pass only.
  This tree also runs the matmul shader pass for f32 coopmat inputs
  (e0bee17), which would otherwise build fp32 q8_1 variants that no
  pipeline creates

Upstream e85e15c (ggml-org#29409) needs no change here: its guards cover the
Intel FA decode shaders this tree does not have, and the cm1 shaders
are only generated when glslc supports coopmat.
pratiknarola-t added a commit to tetherto/qvac-fabric-llm.cpp that referenced this pull request Oct 5, 2026
Ports upstream 70c4e15 (vulkan: int8 coopmat1 matmul implementation
for AMD RDNA3 and RDNA4, ggml-org#27952): the mul_mmq_cm1 shaders, AMD_RDNA4
detection, cm1 int8 shared-memory checks and warptiles, and the q8_1
path selection on coopmat_int_support.

The shaders are byte-identical to upstream, including the hoisted
MUL_MAT_ID row ids and push constants (90c26fc is in this tree). The
pipelines are created the way upstream does it, through
create_mm_pipelines with the tc_mmq_cm1_int and tc_mmq_cm1_int_k tile
tables.

Adapted to this tree:
- quantize_y also requires that src0 needs no reformat copy and that
  src1 is dim01-contiguous. This tree sends non-dim01-contiguous
  quantized src0 through ggml_vk_mul_mat_q_f16 for its quant->f16 copy,
  and ggml_is_contiguous accepts permuted size-1 dims, so without the
  old !y_non_contig term the q8_1 path met the reformat copies (wrong
  MUL_MAT results, and a dispatch of a never-requested copy pipeline).
  The src0 check also fixes int-dot devices without coopmat, where the
  parent already paired the q8_1 path with the src0 copy.
  test_mul_mat_permuted_src0 covers the src0 side.
- AMD_RDNA4 goes into the architecture enum in ggml-vulkan.cpp, since
  ggml-vulkan-types.h (f172be7) is not here
- the IQ4_NL and NVFP4 cases upstream adds to
  ggml_vk_matmul_int_shmem_support are left out: this tree has no
  int-dot q8_1 pipelines for those types, so they would be dead
- the cm1 q8_1 shaders get the device_prefix that the other matmul
  shaders in vulkan-shaders-gen use
- the cm1 q8_1 shaders are generated in the fp16 coopmat pass only.
  This tree also runs the matmul shader pass for f32 coopmat inputs
  (e0bee17), which would otherwise build fp32 q8_1 variants that no
  pipeline creates

Upstream e85e15c (ggml-org#29409) needs no change here: its guards cover the
Intel FA decode shaders this tree does not have, and the cm1 shaders
are only generated when glslc supports coopmat.
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…gml-org#27952)

* vulkan: add int8 coopmat quantized matmul shader

* apply scales inline

* use scalar sums

* probe and directly access coopmat values instead of going through shmem

* add q8_0 support

* add BK_STEP to shader, default to 2

* use larger workgroups

* double buffering

* preload scales

* coopmat load first, then wmma

* use float for scales

* add faster RDNA int->float conversion

* workgroup scheduling for cache proximity

* clean up

* use wave32

* restructure for vgpr use

* skip computation for inactive tiles

* only force subgroup size 32 on AMD RDNA

* use BK_STEP 4

* fix compilation

* move quant-specific prefetch function out of main file

* add q4_1, q5_0, q5_1 support

* restructure mmq cm1 functions

* enable mul_mat_id support

* fix segfault

* fix mul_mat_id bug

* support iq4_nl and mxfp4

* remove elem row/col fast path, invalid for RDNA4

* use shmem arrays for LUTs

* use 4-byte loads where possible

* add q3_k, q4_k, q5_k, q6_k and nvfp4 support

* fix l warptile

* improve performance

* improve performance

* improvements

* dedup b scales

* merge shmem arrays

* undo uint8_t, gate to RDNA3/4

* add RDNA4 architecture, use for hardcoded coopmat elem thread access, set BK_STEP back to 4

* improve offset application

* clean up

* fix iq4_nl and nvfp4 performance

* rdna4 tuning

* use BK_STEP 2 on MUL_MAT_ID

* adapt to upstream changes

* fix shmem support function, clean up comments

* fix warptile logic

Co-authored-by: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com>

* vulkan: add IQ4_XS support to the coopmat1 integer matmul shader (ggml-org#28440)

Adds IQ4_XS to mul_mmq_cm1: dedicated block_a_load/block_a_to_shmem that
expand both nibbles of each packed32 word through cm1_kvalues, LOAD_VEC_A 8
and an IQ4_XS-sized a_panel_bytes estimate for the L2-friendly scheduling.

Assisted-by: OpenAI Codex

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* avoid compiling f16 acc shader variants

---------

Co-authored-by: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
…gml-org#27952)

* vulkan: add int8 coopmat quantized matmul shader

* apply scales inline

* use scalar sums

* probe and directly access coopmat values instead of going through shmem

* add q8_0 support

* add BK_STEP to shader, default to 2

* use larger workgroups

* double buffering

* preload scales

* coopmat load first, then wmma

* use float for scales

* add faster RDNA int->float conversion

* workgroup scheduling for cache proximity

* clean up

* use wave32

* restructure for vgpr use

* skip computation for inactive tiles

* only force subgroup size 32 on AMD RDNA

* use BK_STEP 4

* fix compilation

* move quant-specific prefetch function out of main file

* add q4_1, q5_0, q5_1 support

* restructure mmq cm1 functions

* enable mul_mat_id support

* fix segfault

* fix mul_mat_id bug

* support iq4_nl and mxfp4

* remove elem row/col fast path, invalid for RDNA4

* use shmem arrays for LUTs

* use 4-byte loads where possible

* add q3_k, q4_k, q5_k, q6_k and nvfp4 support

* fix l warptile

* improve performance

* improve performance

* improvements

* dedup b scales

* merge shmem arrays

* undo uint8_t, gate to RDNA3/4

* add RDNA4 architecture, use for hardcoded coopmat elem thread access, set BK_STEP back to 4

* improve offset application

* clean up

* fix iq4_nl and nvfp4 performance

* rdna4 tuning

* use BK_STEP 2 on MUL_MAT_ID

* adapt to upstream changes

* fix shmem support function, clean up comments

* fix warptile logic

Co-authored-by: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com>

* vulkan: add IQ4_XS support to the coopmat1 integer matmul shader (ggml-org#28440)

Adds IQ4_XS to mul_mmq_cm1: dedicated block_a_load/block_a_to_shmem that
expand both nibbles of each packed32 word through cm1_kvalues, LOAD_VEC_A 8
and an IQ4_XS-sized a_panel_bytes estimate for the L2-friendly scheduling.

Assisted-by: OpenAI Codex

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* avoid compiling f16 acc shader variants

---------

Co-authored-by: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 70c4e15)
pratiknarola-t added a commit to tetherto/qvac-fabric-llm.cpp that referenced this pull request Oct 8, 2026
Ports upstream 70c4e15 (vulkan: int8 coopmat1 matmul implementation
for AMD RDNA3 and RDNA4, ggml-org#27952): the mul_mmq_cm1 shaders, AMD_RDNA4
detection, cm1 int8 shared-memory checks and warptiles, and the q8_1
path selection on coopmat_int_support.

The shaders are byte-identical to upstream, including the hoisted
MUL_MAT_ID row ids and push constants (90c26fc is in this tree). The
pipelines are created the way upstream does it, through
create_mm_pipelines with the tc_mmq_cm1_int and tc_mmq_cm1_int_k tile
tables.

Adapted to this tree:
- quantize_y also requires that src0 needs no reformat copy and that
  src1 is dim01-contiguous. This tree sends non-dim01-contiguous
  quantized src0 through ggml_vk_mul_mat_q_f16 for its quant->f16 copy,
  and ggml_is_contiguous accepts permuted size-1 dims, so without the
  old !y_non_contig term the q8_1 path met the reformat copies (wrong
  MUL_MAT results, and a dispatch of a never-requested copy pipeline).
  The src0 check also fixes int-dot devices without coopmat, where the
  parent already paired the q8_1 path with the src0 copy.
  test_mul_mat_permuted_src0 covers the src0 side.
- AMD_RDNA4 goes into the architecture enum in ggml-vulkan.cpp, since
  ggml-vulkan-types.h (f172be7) is not here
- the IQ4_NL and NVFP4 cases upstream adds to
  ggml_vk_matmul_int_shmem_support are left out: this tree has no
  int-dot q8_1 pipelines for those types, so they would be dead
- the cm1 q8_1 shaders get the device_prefix that the other matmul
  shaders in vulkan-shaders-gen use
- the cm1 q8_1 shaders are generated in the fp16 coopmat pass only.
  This tree also runs the matmul shader pass for f32 coopmat inputs
  (e0bee17), which would otherwise build fp32 q8_1 variants that no
  pipeline creates

Upstream e85e15c (ggml-org#29409) needs no change here: its guards cover the
Intel FA decode shaders this tree does not have, and the cm1 shaders
are only generated when glslc supports coopmat.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. Vulkan Issues specific to the Vulkan backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.