Repository navigation
Eval bug: generation performance 3x slower on CUDA compared to llama.cpp #67
Description
Activity
- changed the title
[-]Eval bug:[/-][+]Eval bug: generation performance 4x slower on CUDA compared to llama.cpp[/+]on Jun 12, 2026 - changed the title
[-]Eval bug: generation performance 4x slower on CUDA compared to llama.cpp[/-][+]Eval bug: generation performance 3x slower on CUDA compared to llama.cpp[/+]on Jun 12, 2026 Will investigate.
Thanks for the detailed report and the exact
llama-benchcommand.I found the regression. It was in BeeLlama's CUDA FlashAttention vec decode path, which is the path used by your
q5_0/q4_1cache setup duringtg128 @ d60000. The fork had diverged from upstream by compiling the standard vec FA kernel withminBlocksPerSM = 2instead of upstream's1. That tighter launch bound can reduce the register budget for the D=128 decode kernel, which matches your symptom: prefill stays the same because it uses the MMA path, while decode slows down sharply because it uses the vec path.I fixed this in
2a486938e:- standard cache types now use the upstream-style vec FA launch bound
- Turbo/TCQ cache types keep the tighter occupancy target
- the standard V path no longer has the extra sparse skip that diverged from upstream
Please rebuild current
v0.3.2and rerun the same command. I expecttg128 @ d60000to move much closer to your upstream number. If it is still materially slower, post the new logs and I will continue from that data.Unfortunately, that did not fix inference speed, and prefill speed is notably worse
After
git pull, deleting the build folder and re-running commands:user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk q5_0 -ctv q4_1 -t 4 -ngl 99 -fa on -dio on --device CUDA0 ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB): Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB | model | size | params | backend | ngl | threads | n_batch | n_ubatch | type_k | type_v | fa | dev | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: | TCQ decode: context-adaptive V alpha enabled | minimax-m2 230B.A10B Q3_K - Small | 87.20 GiB | 228.69 B | CUDA | 99 | 4 | 1024 | 1024 | q5_0 | q4_1 | 1 | CUDA0 | pp1000 @ d60000 | 810.52 ± 0.84 | | minimax-m2 230B.A10B Q3_K - Small | 87.20 GiB | 228.69 B | CUDA | 99 | 4 | 1024 | 1024 | q5_0 | q4_1 | 1 | CUDA0 | tg128 @ d60000 | 17.66 ± 0.01 | build: 85e22ea0b (10102)On commit 4caa0a4 branch v0.3.2
I made bench with verbose logging on, but found no notable differences between llama and beellama, still maybe it will help:
beellama.log
llama.logDisregard prefill speed, it was me dropping GPU wattage
Still, according to logs I've attached, problem persists (17.65 ± 0.02 vs 38.91 ± 0.07)I also made benches using another model and kv quants:
user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk turbo3_tcq -ctv turbo3_tcq -t 4 -ngl 99 -fa on -dio on --device CUDA0 ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB): Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB | model | size | params | backend | ngl | threads | n_batch | n_ubatch | type_k | type_v | fa | dev | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: | TCQ: encode V alpha=1.0 (context-adaptive decode-time alpha active) TCQ encode: using shared-memory backtrace (8192 bytes/block) TCQ decode: context-adaptive V alpha enabled | minimax-m2 230B.A10B Q3_K - Small | 87.20 GiB | 228.69 B | CUDA | 99 | 4 | 1024 | 1024 | turbo3_tcq | turbo3_tcq | 1 | CUDA0 | pp1000 @ d60000 | 760.35 ± 0.51 | | minimax-m2 230B.A10B Q3_K - Small | 87.20 GiB | 228.69 B | CUDA | 99 | 4 | 1024 | 1024 | turbo3_tcq | turbo3_tcq | 1 | CUDA0 | tg128 @ d60000 | 10.51 ± 0.01 | build: 85e22ea0b (10102) user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk turbo3 -ctv turbo3 -t 4 -ngl 99 -fa on -dio on --device CUDA0 ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB): Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB | model | size | params | backend | ngl | threads | n_batch | n_ubatch | type_k | type_v | fa | dev | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: | TCQ decode: context-adaptive V alpha enabled | minimax-m2 230B.A10B Q3_K - Small | 87.20 GiB | 228.69 B | CUDA | 99 | 4 | 1024 | 1024 | turbo3 | turbo3 | 1 | CUDA0 | pp1000 @ d60000 | 779.54 ± 0.77 | | minimax-m2 230B.A10B Q3_K - Small | 87.20 GiB | 228.69 B | CUDA | 99 | 4 | 1024 | 1024 | turbo3 | turbo3 | 1 | CUDA0 | tg128 @ d60000 | 51.25 ± 0.19 | build: 85e22ea0b (10102) user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-GGUF/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-Q8_0.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk turbo3 -ctv turbo3 -t 4 -ngl 99 -fa on -dio on --device CUDA0 ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB): Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB | model | size | params | backend | ngl | threads | n_batch | n_ubatch | type_k | type_v | fa | dev | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: | TCQ decode: context-adaptive V alpha enabled | qwen35 27B Q8_0 | 27.04 GiB | 27.32 B | CUDA | 99 | 4 | 1024 | 1024 | turbo3 | turbo3 | 1 | CUDA0 | pp1000 @ d60000 | 1979.43 ± 19.29 | | qwen35 27B Q8_0 | 27.04 GiB | 27.32 B | CUDA | 99 | 4 | 1024 | 1024 | turbo3 | turbo3 | 1 | CUDA0 | tg128 @ d60000 | 42.17 ± 0.05 | build: 85e22ea0b (10102) user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-GGUF/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-Q8_0.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk turbo3_tcq -ctv turbo3_tcq -t 4 -ngl 99 -fa on -dio on --device CUDA0 ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB): Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB | model | size | params | backend | ngl | threads | n_batch | n_ubatch | type_k | type_v | fa | dev | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: | TCQ: encode V alpha=1.0 (context-adaptive decode-time alpha active) TCQ encode: using shared-memory backtrace (8192 bytes/block) TCQ decode: context-adaptive V alpha enabled | qwen35 27B Q8_0 | 27.04 GiB | 27.32 B | CUDA | 99 | 4 | 1024 | 1024 | turbo3_tcq | turbo3_tcq | 1 | CUDA0 | pp1000 @ d60000 | 1911.26 ± 18.98 | | qwen35 27B Q8_0 | 27.04 GiB | 27.32 B | CUDA | 99 | 4 | 1024 | 1024 | turbo3_tcq | turbo3_tcq | 1 | CUDA0 | tg128 @ d60000 | 24.79 ± 0.02 | build: 85e22ea0b (10102) user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-GGUF/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-Q8_0.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk q5_0 -ctv q4_1 -t 4 -ngl 99 -fa on -dio on --device CUDA0 ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB): Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB | model | size | params | backend | ngl | threads | n_batch | n_ubatch | type_k | type_v | fa | dev | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: | TCQ decode: context-adaptive V alpha enabled | qwen35 27B Q8_0 | 27.04 GiB | 27.32 B | CUDA | 99 | 4 | 1024 | 1024 | q5_0 | q4_1 | 1 | CUDA0 | pp1000 @ d60000 | 1941.12 ± 14.71 | | qwen35 27B Q8_0 | 27.04 GiB | 27.32 B | CUDA | 99 | 4 | 1024 | 1024 | q5_0 | q4_1 | 1 | CUDA0 | tg128 @ d60000 | 35.02 ± 0.05 | build: 85e22ea0b (10102)Interestingly, turbo3 quant worked nice with my model, while turbo3_tcq was much slower.
I will try to compile beellama without-DGGML_CUDA_FA=ON. I noticed that this option greatly affected speed in past.I will try to compile beellama without -DGGML_CUDA_FA=ON. I noticed that this option greatly affected speed in past.
I also pulled latest commit 57da314
Nope, that changed nothing. Speeds are exactly the sameThe attached logs still show BeeLlama running
build: 85e22ea0b (10102), which predates the2a486938efix.Yeah, I didn't realize I was on the main branch all along. Sorry 😅
That fixed q5_0/q4_1, but turbo3_tcq is still really slow. Is this expected on my configuration?
user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk q5_0 -ctv q4_1 -t 4 -ngl 99 -fa on -dio on --device CUDA0 ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB): Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB | model | size | params | backend | ngl | threads | n_batch | n_ubatch | type_k | type_v | fa | dev | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: | TCQ decode: context-adaptive V alpha enabled | minimax-m2 230B.A10B Q3_K - Small | 87.20 GiB | 228.69 B | CUDA | 99 | 4 | 1024 | 1024 | q5_0 | q4_1 | 1 | CUDA0 | pp1000 @ d60000 | 811.35 ± 0.55 | | minimax-m2 230B.A10B Q3_K - Small | 87.20 GiB | 228.69 B | CUDA | 99 | 4 | 1024 | 1024 | q5_0 | q4_1 | 1 | CUDA0 | tg128 @ d60000 | 39.40 ± 0.05 | build: 57da314b5 (10264) user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk turbo3_tcq -ctv turbo3_tcq -t 4 -ngl 99 -fa on -dio on --device CUDA0 ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB): Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB | model | size | params | backend | ngl | threads | n_batch | n_ubatch | type_k | type_v | fa | dev | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: | TCQ: encode V alpha=1.0 (context-adaptive decode-time alpha active) TCQ encode: using shared-memory backtrace (8192 bytes/block) TCQ decode: context-adaptive V alpha enabled | minimax-m2 230B.A10B Q3_K - Small | 87.20 GiB | 228.69 B | CUDA | 99 | 4 | 1024 | 1024 | turbo3_tcq | turbo3_tcq | 1 | CUDA0 | pp1000 @ d60000 | 751.92 ± 0.87 | | minimax-m2 230B.A10B Q3_K - Small | 87.20 GiB | 228.69 B | CUDA | 99 | 4 | 1024 | 1024 | turbo3_tcq | turbo3_tcq | 1 | CUDA0 | tg128 @ d60000 | 10.52 ± 0.01 | build: 57da314b5 (10264) user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk turbo3 -ctv turbo3 -t 4 -ngl 99 -fa on -dio on --device CUDA0 ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB): Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB | model | size | params | backend | ngl | threads | n_batch | n_ubatch | type_k | type_v | fa | dev | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: | TCQ decode: context-adaptive V alpha enabled | minimax-m2 230B.A10B Q3_K - Small | 87.20 GiB | 228.69 B | CUDA | 99 | 4 | 1024 | 1024 | turbo3 | turbo3 | 1 | CUDA0 | pp1000 @ d60000 | 775.44 ± 0.41 | | minimax-m2 230B.A10B Q3_K - Small | 87.20 GiB | 228.69 B | CUDA | 99 | 4 | 1024 | 1024 | turbo3 | turbo3 | 1 | CUDA0 | tg128 @ d60000 | 51.57 ± 0.07 | build: 57da314b5 (10264)You know why turbo3 is so fast? It generates garbage :)
user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-cli -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -c 60000 -b 1024 -ub 1024 -ctk turbo3 -ctv turbo3 -t 4 -ngl 99 -fa on --device CUDA0 Loading model... ▄▄ ▄▄ ██ ██ ██ ██ ▀▀█▄ ███▄███▄ ▀▀█▄ ▄████ ████▄ ████▄ ██ ██ ▄█▀██ ██ ██ ██ ▄█▀██ ██ ██ ██ ██ ██ ██ ██ ▀█▄██ ██ ██ ██ ▀█▄██ ██ ▀████ ████▀ ████▀ ██ ██ ▀▀ ▀▀ build : b10264-57da314b5 model : MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf modalities : text available commands: /exit or Ctrl+C stop or exit /regen regenerate the last response /clear clear the chat history /read <file> add a text file /glob <pattern> add text files using globbing pattern > Hello! |TCQ decode: context-adaptive V alpha enabled [Start thinking] ```c,, for a. The扩充, a. 1 (1 云 , so that the For the the the a [1, be THIS - but with the sequence x = re率达到- we The text The "unq5_0/q4_1 are okay
user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-cli -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -c 60000 -b 1024 -ub 1024 -ctk q5_0 -ctv q4_1 -t 4 -ngl 99 -fa on --device CUDA0 Loading model... /TCQ decode: context-adaptive V alpha enabled ▄▄ ▄▄ ██ ██ ██ ██ ▀▀█▄ ███▄███▄ ▀▀█▄ ▄████ ████▄ ████▄ ██ ██ ▄█▀██ ██ ██ ██ ▄█▀██ ██ ██ ██ ██ ██ ██ ██ ▀█▄██ ██ ██ ██ ▀█▄██ ██ ▀████ ████▀ ████▀ ██ ██ ▀▀ ▀▀ build : b10264-57da314b5 model : MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf modalities : text available commands: /exit or Ctrl+C stop or exit /regen regenerate the last response /clear clear the chat history /read <file> add a text file /glob <pattern> add text files using globbing pattern > Hello! [Start thinking] The user says "Hello!" which is a greeting. According to the policy, I should respond politely, greet back, and perhaps ask how I can help. There's no disallowed content. There's no need for a content filter. I should comply. So respond with a friendly greeting, perhaps ask if they need help. So answer accordingly. [End thinking]turbo3_tcq, same garbage
user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-cli -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -c 60000 -b 1024 -ub 1024 -ctk turbo3_tcq -ctv turbo3_tcq -t 4 -ngl 99 -fa on --device CUDA0 Loading model... -TCQ: encode V alpha=1.0 (context-adaptive decode-time alpha active) TCQ encode: using shared-memory backtrace (8192 bytes/block) TCQ decode: context-adaptive V alpha enabled ▄▄ ▄▄ ██ ██ ██ ██ ▀▀█▄ ███▄███▄ ▀▀█▄ ▄████ ████▄ ████▄ ██ ██ ▄█▀██ ██ ██ ██ ▄█▀██ ██ ██ ██ ██ ██ ██ ██ ▀█▄██ ██ ██ ██ ▀█▄██ ██ ▀████ ████▀ ████▀ ██ ██ ▀▀ ▀▀ build : b10264-57da314b5 model : MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf modalities : text available commands: /exit or Ctrl+C stop or exit /regen regenerate the last response /clear clear the chat history /read <file> add a text file /glob <pattern> add text files using globbing pattern > Hello! [Start thinking] The user wants introduce [End thinking] The useruser has br __ [ Prompt: 322.5 t/s | Generation: 99.8 t/s ] > Hello! [Start thinking] The user greeted . The user greet用户 repetitions. The user user"G been you,ولى [End thinking] </think> n="Hello! come</think> <think> 9. Let 1.,éron主人你好</think> :br </think> , ولى 0 and we now,烟<think> </think> </think>, Form </think> nShould I create another issue for this?
It appears that trubo/TCQ types currently are not wired up correctly for models with architecture like MiniMax (D=128), I'll try to fix that. No need to create a separate issue, this should be narrow enough to resolve it there.
In the meanwhile, as you are on v0.3.2 now, you may find the new KVarN cache types to be quite interesting, like kvarn5/4, kvarn4/4, kvarn4/3. They are still "in beta" right now, but I already fixed a lot of issues so it should be more or less stable now, and from my benchmarks they give better precision that other types per bit. Well, if they work with MiniMax. :) But if not, that would be one more report to fix, which is a good thing.
Wow! Well, this is not top-performant choice, but It works:
user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-cli -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -c 60000 -b 1024 -ub 1024 -ctk kvarn4 -ctv kvarn3 -t 4 -ngl 99 -fa on --device CUDA0 Loading model... /TCQ decode: context-adaptive V alpha enabled ▄▄ ▄▄ ██ ██ ██ ██ ▀▀█▄ ███▄███▄ ▀▀█▄ ▄████ ████▄ ████▄ ██ ██ ▄█▀██ ██ ██ ██ ▄█▀██ ██ ██ ██ ██ ██ ██ ██ ▀█▄██ ██ ██ ██ ▀█▄██ ██ ▀████ ████▀ ████▀ ██ ██ ▀▀ ▀▀ build : b10264-57da314b5 model : MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf modalities : text available commands: /exit or Ctrl+C stop or exit /regen regenerate the last response /clear clear the chat history /read <file> add a text file /glob <pattern> add text files using globbing pattern > Hello! [Start thinking] The user said "Hello!". This is a greeting. I should respond in a friendly and welcoming manner, introducing myself and offering to help. [End thinking] Hello! 👋 Welcome! I'm here to help you with any questions you have or tasks you'd like assistance with. How can I help you today? [ Prompt: 251.8 t/s | Generation: 126.9 t/s ] > > /exit Exiting... user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk kvarn4 -ctv kvarn3 -t 4 -ngl 99 -fa on -dio on --device CUDA0 ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB): Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB | model | size | params | backend | ngl | threads | n_batch | n_ubatch | type_k | type_v | fa | dev | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: | TCQ decode: context-adaptive V alpha enabled | minimax-m2 230B.A10B Q3_K - Small | 87.20 GiB | 228.69 B | CUDA | 99 | 4 | 1024 | 1024 | kvarn4 | kvarn3 | 1 | CUDA0 | pp1000 @ d60000 | 798.44 ± 0.53 | | minimax-m2 230B.A10B Q3_K - Small | 87.20 GiB | 228.69 B | CUDA | 99 | 4 | 1024 | 1024 | kvarn4 | kvarn3 | 1 | CUDA0 | tg128 @ d60000 | 26.93 ± 0.05 | build: 57da314b5 (10264) user@faex:~/beellama.cpp$ GGML_VK_VISIBLE_DEVICES=-1 ~/beellama.cpp/build/bin/llama-bench -m /apps/models/unsloth/MiniMax-M2.7-GGUF/UD-Q3_K_S/MiniMax-M2.7-UD-Q3_K_S-00001-of-00003.gguf -p 1000 -d 60000 -b 1024 -ub 1024 -ctk kvarn5 -ctv kvarn4 -t 4 -ngl 99 -fa on -dio on --device CUDA0 ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97288 MiB): Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97288 MiB | model | size | params | backend | ngl | threads | n_batch | n_ubatch | type_k | type_v | fa | dev | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | --: | ------: | ------: | -------: | -----: | -----: | --: | ------------ | --------------: | -------------------: | TCQ decode: context-adaptive V alpha enabled | minimax-m2 230B.A10B Q3_K - Small | 87.20 GiB | 228.69 B | CUDA | 99 | 4 | 1024 | 1024 | kvarn5 | kvarn4 | 1 | CUDA0 | pp1000 @ d60000 | 790.10 ± 0.34 | kvarn4| minimax-m2 230B.A10B Q3_K - Small | 87.20 GiB | 228.69 B | CUDA | 99 | 4 | 1024 | 1024 | kvarn5 | kvarn4 | 1 | CUDA0 | tg128 @ d60000 | 29.58 ± 0.04 | build: 57da314b5 (10264)And compared to q5_0/q4_1:
kvarn5/kvarn4 -- about 1 gb vram saving
kvarn4/kvarn3 -- about 2 gb vram savingOn analytical chat task there is no subjective quality loss. I will try the most radical configuration, kvarn4/kvarn3, on coding tasks later. Thank you for all hard work!
Well, on kvarn5/4 I got sort of more structured output than kvarn4/3, but kvarn5/4 does not work as cache, I got error with every request:
7.02.805.793 W slot update_slots: id 1 | task 14157 | forcing full prompt re-processing because memory cannot trim cached suffix from 5626 (target = 0, draft = 1)There is no such problem on q5_0/q4_1 , it reuses cache nicely on same beellama build
I'm testing kvarn4/kvarn4 with dense Qwen 3.6 27B model. Everything works, the speed it the same as for original MTP model.
I've checked for forcing prompt preprocessing messages and found only single message:34.00.266.612 W slot update_slots: id 0 | task 8034 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see ggml-org#13194 (comment))
Possibly, no issue.
Test plan:
- qwen3.6 27b (unsloth)
- https://qwen.ai/qwencode
- send "1+1" (extremely slow on one rtx5090 without offloading)
- send "2+2" (repeat if have fast response)
- send "1+1"
2.13.485.485 W slot update_slots: id 1 | task 357 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055) 2.22.828.911 W slot update_slots: id 3 | task 358 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)You need 2-3 repeat for reproduce (1+1 -> 2+2 -> 1+1 ...)
Example
docker run -p 11436:11436 --name primary --privileged --rm --gpus all -v /root/.cache/llama.cpp/:/models ghcr.io/anbeeld/beellama.cpp:server-cuda13-preview-v0.3.2 -m "/models/models--unsloth--Qwen3.6-27B-MTP-GGUF/snapshots/5cb35eb3dcbf52dbce5f87dbc64df6aaffadcace/Qwen3.6-27B-UD-Q5_K_XL.gguf" --port 11436 -ctk kvarn4 -ctv kvarn4 --ctx-size 100000 -ngl 99 --temp 0.6 --top-k 20 --top-p 0.95 --min-p 0.00 --spec-type draft-mtp,ngram-mod --jinja --swa-full --ctx-checkpoints 64You can remove
--swa-full&--ctx-checkpoints+ change to-ctk kvarn8 -ctv kvarn8this will not solve the problem
It seems that this problem exists in all llama.cpp and all forksggml-org#22746 (comment)
ggml-org#24055
ggml-org#24176
ggml-org#24785
ggml-org#24797The only solution is vllm
The only solution is vllm
Regarding prefill, it may just have been fixed on llama.cpp
Starting with v0.4.0, BeeLlama uses upstream implementation for DFlash, with some adjustments on top of it.
TurboQuant is not supported anymore as well.
Closing this issue as it's relevant only for pre-v0.4.0.Reacted by Alexander
Name and Version
$ ./build/bin/llama-cli --version
version: 10102 (85e22ea)
built with GNU 15.2.0 for Linux x86_64
Operating systems
Linux
GGML backends
CUDA
Hardware
AMD Ryzen AI Max+ 395
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
Models
https://huggingface.co/unsloth/MiniMax-M2.7-GGUF/tree/main/UD-Q3_K_S
Problem description & steps to reproduce
Use CUDA 13.3 and ubuntu 26.04
Clone upstream llama.cpp and beellama.cpp
Compile both using:
Then benchmark compiled versions (using parameters below)
Prefill speed is mostly the same, but generation speed is about 3 times slower on beellama.cpp on same parameters. Why? What should I tweak?
First Bad Commit
No response
Relevant log output