Skip to content

Eval bug: generation performance 3x slower on CUDA compared to llama.cpp #67

Description

@sakharovaan

Name and Version

$ ./build/bin/llama-cli --version
version: 10102 (85e22ea)
built with GNU 15.2.0 for Linux x86_64

Operating systems

Linux

GGML backends

CUDA

Hardware

AMD Ryzen AI Max+ 395
NVIDIA RTX PRO 6000 Blackwell Workstation Edition

Models

https://huggingface.co/unsloth/MiniMax-M2.7-GGUF/tree/main/UD-Q3_K_S

Problem description & steps to reproduce

Use CUDA 13.3 and ubuntu 26.04
Clone upstream llama.cpp and beellama.cpp

Compile both using:

cmake -B build -DGGML_CUDA=ON -DGGML_NATIVE=ON -DGGML_CUDA_FA_ALL_QUANTS=ON -DCMAKE_BUILD_TYPE=Release -DLLAMA_BUILD_UI=OFF -DLLAMA_BUILD_WEBUI=OFF -DCMAKE_CUDA_ARCHITECTURES=120 -DCUDA_PATH=/usr/local/cuda-13.3 -DGGML_CUDA_FA=ON
cmake --build build -j

Then benchmark compiled versions (using parameters below)

Prefill speed is mostly the same, but generation speed is about 3 times slower on beellama.cpp on same parameters. Why? What should I tweak?

First Bad Commit

No response

Relevant log output

Logs
user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/llama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024 -ub
 1024 -ctk q5_0 -ctv q4_1 -t 4 -ngl 99 -fa on -dio on --device CUDA0
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB):
  Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB
ggml_vulkan: Invalid device index 18446744073709551615 in GGML_VK_VISIBLE_DEVICES.
| model                          |       size |     params | backend    | ngl | threads | n_batch | n_ubatch | type_k | type_v |  fa | dev          |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: |
| minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 |   q5_0 |   q4_1 |   1 | CUDA0        | pp1000 @ d60000 |        984.10 ± 0.94 |
| minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 |   q5_0 |   q4_1 |   1 | CUDA0        |  tg128 @ d60000 |         47.38 ± 0.05 |

build: ebc10770a (9614)
user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024
-ub 1024 -ctk q5_0 -ctv q4_1 -t 4 -ngl 99 -fa on -dio on --device CUDA0
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB):
  Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB
| model                          |       size |     params | backend    | ngl | threads | n_batch | n_ubatch | type_k | type_v |  fa | dev          |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: |
TCQ decode: context-adaptive V alpha enabled
| minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 |   q5_0 |   q4_1 |   1 | CUDA0        | pp1000 @ d60000 |        990.71 ± 0.83 |
| minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 |   q5_0 |   q4_1 |   1 | CUDA0        |  tg128 @ d60000 |         17.65 ± 0.01 |

build: 85e22ea0b (10102)

Activity

  1. changed the title [-]Eval bug:[/-] [+]Eval bug: generation performance 4x slower on CUDA compared to llama.cpp[/+] on Jun 12, 2026
  2. changed the title [-]Eval bug: generation performance 4x slower on CUDA compared to llama.cpp[/-] [+]Eval bug: generation performance 3x slower on CUDA compared to llama.cpp[/+] on Jun 12, 2026
  3. Anbeeld commented on Jun 12, 2026

    @Anbeeld
    Owner

    Will investigate.

  4. Anbeeld commented on Jun 15, 2026

    @Anbeeld
    Owner

    Thanks for the detailed report and the exact llama-bench command.

    I found the regression. It was in BeeLlama's CUDA FlashAttention vec decode path, which is the path used by your q5_0/q4_1 cache setup during tg128 @ d60000. The fork had diverged from upstream by compiling the standard vec FA kernel with minBlocksPerSM = 2 instead of upstream's 1. That tighter launch bound can reduce the register budget for the D=128 decode kernel, which matches your symptom: prefill stays the same because it uses the MMA path, while decode slows down sharply because it uses the vec path.

    I fixed this in 2a486938e:

    • standard cache types now use the upstream-style vec FA launch bound
    • Turbo/TCQ cache types keep the tighter occupancy target
    • the standard V path no longer has the extra sparse skip that diverged from upstream

    Please rebuild current v0.3.2 and rerun the same command. I expect tg128 @ d60000 to move much closer to your upstream number. If it is still materially slower, post the new logs and I will continue from that data.

  5. sakharovaan commented on Jun 15, 2026

    @sakharovaan
    Author

    Unfortunately, that did not fix inference speed, and prefill speed is notably worse

    After git pull, deleting the build folder and re-running commands:

    user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk q5_0 -ctv q4_1 -t 4 -ngl 99 -fa on -dio on --device CUDA0
    ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB):
      Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB
    | model                          |       size |     params | backend    | ngl | threads | n_batch | n_ubatch | type_k | type_v |  fa | dev          |            test |                  t/s |
    | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: |
    TCQ decode: context-adaptive V alpha enabled
    | minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 |   q5_0 |   q4_1 |   1 | CUDA0        | pp1000 @ d60000 |        810.52 ± 0.84 |
    
    
    | minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 |   q5_0 |   q4_1 |   1 | CUDA0        |  tg128 @ d60000 |         17.66 ± 0.01 |
    
    build: 85e22ea0b (10102)
    

    On commit 4caa0a4 branch v0.3.2

    I made bench with verbose logging on, but found no notable differences between llama and beellama, still maybe it will help:
    beellama.log
    llama.log

  6. sakharovaan commented on Jun 15, 2026

    @sakharovaan
    Author

    Disregard prefill speed, it was me dropping GPU wattage
    Still, according to logs I've attached, problem persists (17.65 ± 0.02 vs 38.91 ± 0.07)

  7. sakharovaan commented on Jun 15, 2026

    @sakharovaan
    Author

    I also made benches using another model and kv quants:

    user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024
    -ub 1024 -ctk turbo3_tcq -ctv turbo3_tcq -t 4 -ngl 99 -fa on -dio on --device CUDA0
    ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB):
      Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB
    | model                          |       size |     params | backend    | ngl | threads | n_batch | n_ubatch | type_k | type_v |  fa | dev          |            test |                  t/s |
    | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: |
    TCQ: encode V alpha=1.0 (context-adaptive decode-time alpha active)
    TCQ encode: using shared-memory backtrace (8192 bytes/block)
    TCQ decode: context-adaptive V alpha enabled
    | minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 | turbo3_tcq | turbo3_tcq |   1 | CUDA0        | pp1000 @ d60000 |        760.35 ± 0.51 |
    | minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 | turbo3_tcq | turbo3_tcq |   1 | CUDA0        |  tg128 @ d60000 |         10.51 ± 0.01 |
    
    build: 85e22ea0b (10102)
    user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk turbo3 -ctv turbo3 -t 4 -ngl 99 -fa on -dio on --device CUDA0
    ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB):
      Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB
    | model                          |       size |     params | backend    | ngl | threads | n_batch | n_ubatch | type_k | type_v |  fa | dev          |            test |                  t/s |
    | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: |
    TCQ decode: context-adaptive V alpha enabled
    | minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 | turbo3 | turbo3 |   1 | CUDA0        | pp1000 @ d60000 |        779.54 ± 0.77 |
    | minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 | turbo3 | turbo3 |   1 | CUDA0        |  tg128 @ d60000 |         51.25 ± 0.19 |
    
    build: 85e22ea0b (10102)
    user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-GGUF/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-Q8_0.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk turbo3 -ctv turbo3 -t 4 -ngl 99 -fa on -dio on --device CUDA0
    ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB):
      Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB
    | model                          |       size |     params | backend    | ngl | threads | n_batch | n_ubatch | type_k | type_v |  fa | dev          |            test |                  t/s |
    | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: |
    TCQ decode: context-adaptive V alpha enabled
    | qwen35 27B Q8_0                |  27.04 GiB |    27.32 B | CUDA       |  99 |       4 |    1024 |     1024 | turbo3 | turbo3 |   1 | CUDA0        | pp1000 @ d60000 |      1979.43 ± 19.29 |
    | qwen35 27B Q8_0                |  27.04 GiB |    27.32 B | CUDA       |  99 |       4 |    1024 |     1024 | turbo3 | turbo3 |   1 | CUDA0        |  tg128 @ d60000 |         42.17 ± 0.05 |
    
    build: 85e22ea0b (10102)
    user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-GGUF/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-Q8_0.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk turbo3_tcq -ctv turbo3_tcq -t 4 -ngl 99 -fa on -dio on --device CUDA0
    ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB):
      Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB
    | model                          |       size |     params | backend    | ngl | threads | n_batch | n_ubatch | type_k | type_v |  fa | dev          |            test |                  t/s |
    | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: |
    TCQ: encode V alpha=1.0 (context-adaptive decode-time alpha active)
    TCQ encode: using shared-memory backtrace (8192 bytes/block)
    TCQ decode: context-adaptive V alpha enabled
    | qwen35 27B Q8_0                |  27.04 GiB |    27.32 B | CUDA       |  99 |       4 |    1024 |     1024 | turbo3_tcq | turbo3_tcq |   1 | CUDA0        | pp1000 @ d60000 |      1911.26 ± 18.98 |
    | qwen35 27B Q8_0                |  27.04 GiB |    27.32 B | CUDA       |  99 |       4 |    1024 |     1024 | turbo3_tcq | turbo3_tcq |   1 | CUDA0        |  tg128 @ d60000 |         24.79 ± 0.02 |
    
    build: 85e22ea0b (10102)
    user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-GGUF/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-Q8_0.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk q5_0 -ctv q4_1 -t 4 -ngl 99 -fa on -dio on --device CUDA0
    ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB):
      Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB
    | model                          |       size |     params | backend    | ngl | threads | n_batch | n_ubatch | type_k | type_v |  fa | dev          |            test |                  t/s |
    | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: |
    TCQ decode: context-adaptive V alpha enabled
    | qwen35 27B Q8_0                |  27.04 GiB |    27.32 B | CUDA       |  99 |       4 |    1024 |     1024 |   q5_0 |   q4_1 |   1 | CUDA0        | pp1000 @ d60000 |      1941.12 ± 14.71 |
    | qwen35 27B Q8_0                |  27.04 GiB |    27.32 B | CUDA       |  99 |       4 |    1024 |     1024 |   q5_0 |   q4_1 |   1 | CUDA0        |  tg128 @ d60000 |         35.02 ± 0.05 |
    
    build: 85e22ea0b (10102)
    

    Interestingly, turbo3 quant worked nice with my model, while turbo3_tcq was much slower.
    I will try to compile beellama without -DGGML_CUDA_FA=ON. I noticed that this option greatly affected speed in past.

  8. sakharovaan commented on Jun 15, 2026

    @sakharovaan
    Author

    I will try to compile beellama without -DGGML_CUDA_FA=ON. I noticed that this option greatly affected speed in past.

    I also pulled latest commit 57da314
    Nope, that changed nothing. Speeds are exactly the same

  9. Anbeeld commented on Jun 15, 2026

    @Anbeeld
    Owner

    The attached logs still show BeeLlama running build: 85e22ea0b (10102), which predates the 2a486938e fix.

  10. sakharovaan commented on Jun 15, 2026

    @sakharovaan
    Author

    Yeah, I didn't realize I was on the main branch all along. Sorry 😅

    That fixed q5_0/q4_1, but turbo3_tcq is still really slow. Is this expected on my configuration?

    user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk q5_0 -ctv q4_1 -t 4 -ngl 99 -fa on -dio on --device CUDA0
    ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB):
      Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB
    | model                          |       size |     params | backend    | ngl | threads | n_batch | n_ubatch | type_k | type_v |  fa | dev          |            test |                  t/s |
    | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: |
    TCQ decode: context-adaptive V alpha enabled
    | minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 |   q5_0 |   q4_1 |   1 | CUDA0        | pp1000 @ d60000 |        811.35 ± 0.55 |
    | minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 |   q5_0 |   q4_1 |   1 | CUDA0        |  tg128 @ d60000 |         39.40 ± 0.05 |
    
    build: 57da314b5 (10264)
    user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk turbo3_tcq -ctv turbo3_tcq -t 4 -ngl 99 -fa on -dio on --device CUDA0
    ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB):
      Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB
    | model                          |       size |     params | backend    | ngl | threads | n_batch | n_ubatch | type_k | type_v |  fa | dev          |            test |                  t/s |
    | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: |
    TCQ: encode V alpha=1.0 (context-adaptive decode-time alpha active)
    TCQ encode: using shared-memory backtrace (8192 bytes/block)
    TCQ decode: context-adaptive V alpha enabled
    | minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 | turbo3_tcq | turbo3_tcq |   1 | CUDA0        | pp1000 @ d60000 |        751.92 ± 0.87 |
    | minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 | turbo3_tcq | turbo3_tcq |   1 | CUDA0        |  tg128 @ d60000 |         10.52 ± 0.01 |
    
    build: 57da314b5 (10264)
    user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk turbo3 -ctv turbo3 -t 4 -ngl 99 -fa on -dio on --device CUDA0
    ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB):
      Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB
    | model                          |       size |     params | backend    | ngl | threads | n_batch | n_ubatch | type_k | type_v |  fa | dev          |            test |                  t/s |
    | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: |
    TCQ decode: context-adaptive V alpha enabled
    | minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 | turbo3 | turbo3 |   1 | CUDA0        | pp1000 @ d60000 |        775.44 ± 0.41 |
    | minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 | turbo3 | turbo3 |   1 | CUDA0        |  tg128 @ d60000 |         51.57 ± 0.07 |
    
    build: 57da314b5 (10264)
    
  11. sakharovaan commented on Jun 15, 2026

    @sakharovaan
    Author

    You know why turbo3 is so fast? It generates garbage :)

    user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-cli -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -c 60000 -b 1024 -ub 1024 -ctk turbo3 -ctv turbo3 -t 4 -ngl 99 -fa on --device CUDA0
    
    Loading model...
    
    
    ▄▄ ▄▄
    ██ ██
    ██ ██  ▀▀█▄ ███▄███▄  ▀▀█▄    ▄████ ████▄ ████▄
    ██ ██ ▄█▀██ ██ ██ ██ ▄█▀██    ██    ██ ██ ██ ██
    ██ ██ ▀█▄██ ██ ██ ██ ▀█▄██ ██ ▀████ ████▀ ████▀
                                        ██    ██
                                        ▀▀    ▀▀
    
    build      : b10264-57da314b5
    model      : MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf
    modalities : text
    
    available commands:
      /exit or Ctrl+C     stop or exit
      /regen              regenerate the last response
      /clear              clear the chat history
      /read <file>        add a text file
      /glob <pattern>     add text files using globbing pattern
    
    
    > Hello!
    
    |TCQ decode: context-adaptive V alpha enabled
    [Start thinking]
    ```c,, for a.
    
    The扩充, a. 1 (1
    云
    , so that the
    
    For the the the a
    
    [1, be
    
    THIS - but with the sequence x = re率达到- we
    
    The text
    
    The "un
    

    q5_0/q4_1 are okay

    user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-cli -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -c 60000 -b 1024 -ub 1024 -ctk q5_0 -ctv q4_1 -t 4 -ngl 99 -fa on --device CUDA0
    
    Loading model... /TCQ decode: context-adaptive V alpha enabled
    
    
    
    ▄▄ ▄▄
    ██ ██
    ██ ██  ▀▀█▄ ███▄███▄  ▀▀█▄    ▄████ ████▄ ████▄
    ██ ██ ▄█▀██ ██ ██ ██ ▄█▀██    ██    ██ ██ ██ ██
    ██ ██ ▀█▄██ ██ ██ ██ ▀█▄██ ██ ▀████ ████▀ ████▀
                                        ██    ██
                                        ▀▀    ▀▀
    
    build      : b10264-57da314b5
    model      : MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf
    modalities : text
    
    available commands:
      /exit or Ctrl+C     stop or exit
      /regen              regenerate the last response
      /clear              clear the chat history
      /read <file>        add a text file
      /glob <pattern>     add text files using globbing pattern
    
    
    > Hello!
    
    [Start thinking]
    The user says "Hello!" which is a greeting. According to the policy, I should respond politely, greet back, and perhaps ask how I can help. There's no disallowed content. There's no need for a content filter. I should comply. So respond with a friendly greeting, perhaps ask if they need help. So answer accordingly.
    
    [End thinking]
    

    turbo3_tcq, same garbage

    user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-cli -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -c 60000 -b 1024 -ub 1024 -ctk turbo3_tcq -ctv turbo3_tcq -t 4 -ngl 99 -fa on --device CUDA0
    
    Loading model... -TCQ: encode V alpha=1.0 (context-adaptive decode-time alpha active)
    TCQ encode: using shared-memory backtrace (8192 bytes/block)
    TCQ decode: context-adaptive V alpha enabled
    
    
    
    ▄▄ ▄▄
    ██ ██
    ██ ██  ▀▀█▄ ███▄███▄  ▀▀█▄    ▄████ ████▄ ████▄
    ██ ██ ▄█▀██ ██ ██ ██ ▄█▀██    ██    ██ ██ ██ ██
    ██ ██ ▀█▄██ ██ ██ ██ ▀█▄██ ██ ▀████ ████▀ ████▀
                                        ██    ██
                                        ▀▀    ▀▀
    
    build      : b10264-57da314b5
    model      : MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf
    modalities : text
    
    available commands:
      /exit or Ctrl+C     stop or exit
      /regen              regenerate the last response
      /clear              clear the chat history
      /read <file>        add a text file
      /glob <pattern>     add text files using globbing pattern
    
    
    > Hello!
    
    [Start thinking]
    The user wants introduce
    
    
    [End thinking]
    
    The useruser has
    br
    __
    
    
    [ Prompt: 322.5 t/s | Generation: 99.8 t/s ]
    
    > Hello!
    
    [Start thinking]
    The user greeted
    . The user greet用户 repetitions. The user user"G been you,ولى
    [End thinking]
    
    </think>
    n="Hello!
    come</think>
    <think>
    9. Let
    1.,éron主人你好</think>
    :br
    </think>
    , ولى
    
    
    0 and we now,烟<think>
    </think>
    </think>, Form
    </think>
    n
    

    Should I create another issue for this?

  12. Anbeeld commented on Jun 15, 2026

    @Anbeeld
    Owner

    It appears that trubo/TCQ types currently are not wired up correctly for models with architecture like MiniMax (D=128), I'll try to fix that. No need to create a separate issue, this should be narrow enough to resolve it there.

    In the meanwhile, as you are on v0.3.2 now, you may find the new KVarN cache types to be quite interesting, like kvarn5/4, kvarn4/4, kvarn4/3. They are still "in beta" right now, but I already fixed a lot of issues so it should be more or less stable now, and from my benchmarks they give better precision that other types per bit. Well, if they work with MiniMax. :) But if not, that would be one more report to fix, which is a good thing.

  13. sakharovaan commented on Jun 15, 2026

    @sakharovaan
    Author

    Wow! Well, this is not top-performant choice, but It works:

    user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-cli -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -c 60000 -b 1024 -ub 1024 -ctk kvarn4 -ctv kvarn3 -t 4 -ngl 99 -fa on --device CUDA0
    
    Loading model... /TCQ decode: context-adaptive V alpha enabled
    
    
    
    ▄▄ ▄▄
    ██ ██
    ██ ██  ▀▀█▄ ███▄███▄  ▀▀█▄    ▄████ ████▄ ████▄
    ██ ██ ▄█▀██ ██ ██ ██ ▄█▀██    ██    ██ ██ ██ ██
    ██ ██ ▀█▄██ ██ ██ ██ ▀█▄██ ██ ▀████ ████▀ ████▀
                                        ██    ██
                                        ▀▀    ▀▀
    
    build      : b10264-57da314b5
    model      : MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf
    modalities : text
    
    available commands:
      /exit or Ctrl+C     stop or exit
      /regen              regenerate the last response
      /clear              clear the chat history
      /read <file>        add a text file
      /glob <pattern>     add text files using globbing pattern
    
    
    > Hello!
    
    [Start thinking]
    The user said "Hello!". This is a greeting. I should respond in a friendly and welcoming manner, introducing myself and offering to help.
    
    [End thinking]
    
    Hello! 👋 Welcome! I'm here to help you with any questions you have or tasks you'd like assistance with. How can I help you today?
    
    [ Prompt: 251.8 t/s | Generation: 126.9 t/s ]
    
    >
    
    > /exit
    
    
    Exiting...
    user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk kvarn4 -ctv kvarn3 -t 4 -ngl 99 -fa on -dio on --device CUDA0
    ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB):
      Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB
    | model                          |       size |     params | backend    | ngl | threads | n_batch | n_ubatch | type_k | type_v |  fa | dev          |            test |                  t/s |
    | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: |
    TCQ decode: context-adaptive V alpha enabled
    | minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 | kvarn4 | kvarn3 |   1 | CUDA0        | pp1000 @ d60000 |        798.44 ± 0.53 |
    | minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 | kvarn4 | kvarn3 |   1 | CUDA0        |  tg128 @ d60000 |         26.93 ± 0.05 |
    
    build: 57da314b5 (10264)
    user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk kvarn5 -ctv kvarn4 -t 4 -ngl 99 -fa on -dio on --device CUDA0
    ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB):
      Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB
    | model                          |       size |     params | backend    | ngl | threads | n_batch | n_ubatch | type_k | type_v |  fa | dev          |            test |                  t/s |
    | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: |
    TCQ decode: context-adaptive V alpha enabled
    | minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 | kvarn5 | kvarn4 |   1 | CUDA0        | pp1000 @ d60000 |        790.10 ± 0.34 |
    kvarn4| minimax-m2 230B.A10B Q3_K - Small |  87.20 GiB |   228.69 B | CUDA       |  99 |       4 |    1024 |     1024 | kvarn5 | kvarn4 |   1 | CUDA0        |  tg128 @ d60000 |         29.58 ± 0.04 |
    
    build: 57da314b5 (10264)
    

    And compared to q5_0/q4_1:
    kvarn5/kvarn4 -- about 1 gb vram saving
    kvarn4/kvarn3 -- about 2 gb vram saving

    On analytical chat task there is no subjective quality loss. I will try the most radical configuration, kvarn4/kvarn3, on coding tasks later. Thank you for all hard work!

  14. sakharovaan commented on Jun 15, 2026

    @sakharovaan
    Author

    Well, on kvarn5/4 I got sort of more structured output than kvarn4/3, but kvarn5/4 does not work as cache, I got error with every request:

    7.02.805.793 W slot update_slots: id 1 | task 14157 | forcing full prompt re-processing because memory cannot trim cached suffix from 5626 (target = 0, draft = 1)

    There is no such problem on q5_0/q4_1 , it reuses cache nicely on same beellama build

  15. ethernidee commented on Jun 16, 2026

    @ethernidee

    I'm testing kvarn4/kvarn4 with dense Qwen 3.6 27B model. Everything works, the speed it the same as for original MTP model.
    I've checked for forcing prompt preprocessing messages and found only single message:

    34.00.266.612 W slot update_slots: id 0 | task 8034 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see ggml-org#13194 (comment))

    Possibly, no issue.

    log.txt

  16. leagueofsoups commented on Jun 20, 2026

    @leagueofsoups

    Test plan:

    • qwen3.6 27b (unsloth)
    • https://qwen.ai/qwencode
    • send "1+1" (extremely slow on one rtx5090 without offloading)
    • send "2+2" (repeat if have fast response)
    • send "1+1"
    2.13.485.485 W slot update_slots: id  1 | task 357 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
    2.22.828.911 W slot update_slots: id  3 | task 358 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
    

    You need 2-3 repeat for reproduce (1+1 -> 2+2 -> 1+1 ...)

    Example

          docker run -p 11436:11436 --name primary --privileged --rm --gpus all -v /root/.cache/llama.cpp/:/models ghcr.io/anbeeld/beellama.cpp:server-cuda13-preview-v0.3.2
          -m "/models/models--unsloth--Qwen3.6-27B-MTP-GGUF/snapshots/5cb35eb3dcbf52dbce5f87dbc64df6aaffadcace/Qwen3.6-27B-UD-Q5_K_XL.gguf"
          --port 11436
          -ctk kvarn4 -ctv kvarn4
          --ctx-size 100000 -ngl 99 --temp 0.6 --top-k 20 --top-p 0.95 --min-p 0.00
          --spec-type draft-mtp,ngram-mod
          --jinja
          --swa-full
          --ctx-checkpoints 64
    

    You can remove --swa-full & --ctx-checkpoints + change to -ctk kvarn8 -ctv kvarn8 this will not solve the problem
    It seems that this problem exists in all llama.cpp and all forks

    ggml-org#22746 (comment)
    ggml-org#24055
    ggml-org#24176
    ggml-org#24785
    ggml-org#24797

    The only solution is vllm

  17. laurentftech commented on Jun 26, 2026

    @laurentftech

    The only solution is vllm

    Regarding prefill, it may just have been fixed on llama.cpp

    ggml-org#24055 (comment)

  18. Anbeeld commented on Jul 19, 2026

    @Anbeeld
    Owner

    Starting with v0.4.0, BeeLlama uses upstream implementation for DFlash, with some adjustments on top of it.
    TurboQuant is not supported anymore as well.
    Closing this issue as it's relevant only for pre-v0.4.0.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions