Skip to content

Misc. bug: Linux Vulkan build is slower than CPU with Ternary 8B, and 10 times slower than a larger model #28

Description

@pmanna

Name and Version

./llama-cli --version
load_backend: loaded RPC backend from ~/Applications/llama-prism/libggml-rpc.so
load_backend: loaded Vulkan backend from ~/Applications/llama-prism/libggml-vulkan.so
load_backend: loaded CPU backend from ~/Applications/llama-prism/libggml-cpu-zen4.so
version: 8846 (d104cf1b6)
built with GNU 11.4.0 for Linux x86_64
llama-cli --list-devices
load_backend: loaded RPC backend from ~/Applications/llama-prism/libggml-rpc.so
load_backend: loaded Vulkan backend from ~/Applications/llama-prism/libggml-vulkan.so
load_backend: loaded CPU backend from ~/Applications/llama-prism/libggml-cpu-zen4.so
Available devices:
  Vulkan0: AMD Radeon 780M Graphics (RADV PHOENIX) (40361 MiB, 19410 MiB free)

Operating systems

Linux

Which llama.cpp modules do you know to be affected?

llama-server

Command line

./llama-server -m ~/.lmstudio/models/prism-ml/Ternary-Bonsai-8B-gguf/Ternary-Bonsai-8B-Q2_0.gguf \
  --alias bonsai-8b \
  --ctx-size 8192 \
  --jinja --gpu-layers all \
  --temp 0.3 --top-p 1.0 --min-p 0.01 \
  --sleep-idle-seconds 600 \
  --host 0.0.0.0 --port 1234

Problem description & steps to reproduce

I've been trying the most recent Ternary 8B model with the Linux Vulkan build, and the performance is terrible (2.4 t/s), even worse than the CPU build (2.73 t/s). For comparison, the same hardware runs a much larger model (Qwen 3.6 MoE) at 24 t/s. The GPU has 16 GB dedicated, and it only gets filled in at 25% (~4 GB at 8K context). Interestingly enough, when chatting, the GPU is barely involved, while the CPU spikes, even with --gpu-layers 'all'. From my POV, it looks like the Vulkan build "forgets" that it has a GPU available for some tasks, but of course I may be wrong

OS Specs

OS: Linux Mint 22.3 x86_64
Host: Venus series
Kernel: 6.8.0-110-generic
Shell: bash 5.2.21
Terminal: /dev/pts/0
CPU: AMD Ryzen 9 7940HS w/ Radeon 780M Graphics (16) @ 5.263GHz
GPU: AMD ATI c4:00.0 Phoenix1
Memory: 3545MiB / 47954MiB

Activity

  1. changed the title [-]Misc. bug: Ternary Linux Vulkan buld is slower than CPU build, and 10 times slower than a larger model[/-] [+]Misc. bug: Ternary Linux Vulkan build is slower than CPU build, and 10 times slower than a larger model[/+] on Apr 22, 2026
  2. changed the title [-]Misc. bug: Ternary Linux Vulkan build is slower than CPU build, and 10 times slower than a larger model[/-] [+]Misc. bug: Linux Vulkan build is slower than CPU with Ternary 8B, and 10 times slower than a larger model[/+] on Apr 22, 2026
  3. shrisha108 commented on Apr 22, 2026

    @shrisha108

    yes, same here. exactly same situation. Fedora 43. Amd RX 6950XT model loaded into VRAM and Vulkan device found and used, but speed is 5 tokents/s

  4. khosravipasha commented on Apr 22, 2026

    @khosravipasha
    Collaborator

    Ternary Bonsai packed into Q2_0 currently does not have Vulkan backend yet. For Q2_0 we have Metal, CUDA, ARM cpu backends at the moment. Everything else falls back to generic cpu which is slow until we also an optimzied x86 backend.

    The normal Bonsai models (Q1_0) have vulkan backend so that should be fast.

  5. khosravipasha commented on Apr 22, 2026

    @khosravipasha
    Collaborator

    For AMD gpu might be able to get much faster results with the HIP build until vulkan support is added.

    Linux (AMD):
    Ubuntu x64 (ROCm 7.2)

  6. khosravipasha commented on Apr 22, 2026

    @khosravipasha
    Collaborator

    See this one on how to run on AMD (this was with Q1_0) but I think similar approach works for Ternary-Bonsai packed into Q2_0
    https://github.com/PrismML-Eng/Bonsai-demo/blob/main/community-benchmarks/bonsai/rocm-hip-strix-halo-128gb-archlinux.md

  7. pmanna commented on Apr 22, 2026

    @pmanna
    Author

    Thanks @khosravipasha ,

    Unfortunately for me, the ROCm solution doesn't work for my integrated GPU: I'll wait for the Vulkan support to be finalized.

  8. khosravipasha commented on May 11, 2026

    @khosravipasha
    Collaborator

    There is a new PR on vulkan Q2_0, going to review and try to merge in our fork sometime this week, in the meantime can try it if curious #32

  9. JadeTusk commented on May 15, 2026

    @JadeTusk

    I have the same issue on Windows 11, but the ROCm version worked really well for me (discreet GPU).

  10. admorelli commented on May 24, 2026

    @admorelli

    Using up to date CachyOS, for me does not even finish warming up, PS already have the PR #32 .

  11. added a commit that references this issue on Jul 15, 2026
  12. khosravipasha commented on Jul 20, 2026

    @khosravipasha
    Collaborator

    We have a variant of Q2_0 (group size 64) merged in llama.cpp now and alreayd has NEON/Metal/Vulkan merbed, waiting for CUDA, for those need to use Q2_0_g64.gguf file names from our huggingface. See https://github.com/PrismML-Eng/Bonsai-demo#upstream-status-for-ternary

    Closing this for now, let me know if issue persists.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions