Repository navigation
Misc. bug: Linux Vulkan build is slower than CPU with Ternary 8B, and 10 times slower than a larger model #28
Description
Activity
- changed the title
[-]Misc. bug: Ternary Linux Vulkan buld is slower than CPU build, and 10 times slower than a larger model[/-][+]Misc. bug: Ternary Linux Vulkan build is slower than CPU build, and 10 times slower than a larger model[/+]on Apr 22, 2026 - changed the title
[-]Misc. bug: Ternary Linux Vulkan build is slower than CPU build, and 10 times slower than a larger model[/-][+]Misc. bug: Linux Vulkan build is slower than CPU with Ternary 8B, and 10 times slower than a larger model[/+]on Apr 22, 2026 yes, same here. exactly same situation. Fedora 43. Amd RX 6950XT model loaded into VRAM and Vulkan device found and used, but speed is 5 tokents/s
Ternary Bonsai packed into Q2_0 currently does not have Vulkan backend yet. For Q2_0 we have Metal, CUDA, ARM cpu backends at the moment. Everything else falls back to generic cpu which is slow until we also an optimzied x86 backend.
The normal Bonsai models (Q1_0) have vulkan backend so that should be fast.
For AMD gpu might be able to get much faster results with the HIP build until vulkan support is added.
Linux (AMD):
Ubuntu x64 (ROCm 7.2)See this one on how to run on AMD (this was with Q1_0) but I think similar approach works for Ternary-Bonsai packed into Q2_0
https://github.com/PrismML-Eng/Bonsai-demo/blob/main/community-benchmarks/bonsai/rocm-hip-strix-halo-128gb-archlinux.mdThanks @khosravipasha ,
Unfortunately for me, the ROCm solution doesn't work for my integrated GPU: I'll wait for the Vulkan support to be finalized.
- added 7 commits that reference this issue
on May 3, 2026 There is a new PR on vulkan Q2_0, going to review and try to merge in our fork sometime this week, in the meantime can try it if curious #32
Reacted by Paolo MannaI have the same issue on Windows 11, but the ROCm version worked really well for me (discreet GPU).
Reacted by Pasha KhosraviUsing up to date CachyOS, for me does not even finish warming up, PS already have the PR #32 .
We have a variant of Q2_0 (group size 64) merged in llama.cpp now and alreayd has NEON/Metal/Vulkan merbed, waiting for CUDA, for those need to use Q2_0_g64.gguf file names from our huggingface. See https://github.com/PrismML-Eng/Bonsai-demo#upstream-status-for-ternary
Closing this for now, let me know if issue persists.
Name and Version
Operating systems
Linux
Which llama.cpp modules do you know to be affected?
llama-server
Command line
./llama-server -m ~/.lmstudio/models/prism-ml/Ternary-Bonsai-8B-gguf/Ternary-Bonsai-8B-Q2_0.gguf \ --alias bonsai-8b \ --ctx-size 8192 \ --jinja --gpu-layers all \ --temp 0.3 --top-p 1.0 --min-p 0.01 \ --sleep-idle-seconds 600 \ --host 0.0.0.0 --port 1234Problem description & steps to reproduce
I've been trying the most recent Ternary 8B model with the Linux Vulkan build, and the performance is terrible (2.4 t/s), even worse than the CPU build (2.73 t/s). For comparison, the same hardware runs a much larger model (Qwen 3.6 MoE) at 24 t/s. The GPU has 16 GB dedicated, and it only gets filled in at 25% (~4 GB at 8K context). Interestingly enough, when chatting, the GPU is barely involved, while the CPU spikes, even with
--gpu-layers 'all'. From my POV, it looks like the Vulkan build "forgets" that it has a GPU available for some tasks, but of course I may be wrongOS Specs