Skip to content

vulkan: command buffer graph reuse - #24720

Draft
0cc4m wants to merge 6 commits into
masterfrom
0cc4m/vulkan-graph-reuse
Draft

0cc4m wants to merge 6 commits into
masterfrom
0cc4m/vulkan-graph-reuse

Conversation

@0cc4m

@0cc4m 0cc4m commented Jun 17, 2026

Copy link
Copy Markdown
Contributor

Overview

This is more of a prototype to see if reusing graphs/command buffers is possible and what difference it makes. I cannot see much of a performance difference at all from this, which suggests we're already overlaying GPU and CPU work well-enough that not much is to gain from reducing CPU work.

However, my list of devices to test on is limited, and it may also be advantageous to reduce the CPU load in this way, by not rerecording identical graphs each token.

@jeffbolznv Let me know what you think, whether you think this could be worth adding or not (and not necessarily this specific implementation, I didn't think it through enough yet)

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, written by AI, tested and refined by me

@github-actions github-actions Bot added Vulkan Issues specific to the Vulkan backend ggml changes relating to the ggml tensor library for machine learning labels Jun 17, 2026
@0cc4m

0cc4m commented Jun 17, 2026

Copy link
Copy Markdown
Contributor Author

Some benchmarks, all on AMD EPYC 7302 server:

AMD Radeon Pro VII:

model size params ngl fa mmap test t/s (before) t/s (after) diff
llama 8B Q4_0 4.33 GiB 8.03 B -1 1 0 pp512 845.26 ± 4.65 844.53 ± 3.67 -0.1%
llama 8B Q4_0 4.33 GiB 8.03 B -1 1 0 tg128 101.98 ± 1.07 102.17 ± 1.12 +0.2%
llama 1B Q8_0 1.84 GiB 1.86 B -1 1 0 pp512 2952.59 ± 3.03 2946.52 ± 5.29 -0.2%
llama 1B Q8_0 1.84 GiB 1.86 B -1 1 0 tg128 222.78 ± 0.93 222.76 ± 4.41 -0.0%
deepseek2 30B.A3B Q3_K - Small 12.37 GiB 29.94 B -1 1 0 pp512 825.11 ± 2.82 823.26 ± 3.06 -0.2%
deepseek2 30B.A3B Q3_K - Small 12.37 GiB 29.94 B -1 1 0 tg128 71.27 ± 1.20 71.72 ± 0.61 +0.6%

Intel A770:

model size params ngl fa mmap test t/s (before) t/s (after) diff
llama 8B Q4_0 4.33 GiB 8.03 B -1 1 0 pp512 1278.56 ± 5.22 1279.14 ± 3.72 +0.0%
llama 8B Q4_0 4.33 GiB 8.03 B -1 1 0 tg128 42.83 ± 0.05 42.73 ± 0.13 -0.2%
llama 1B Q8_0 1.84 GiB 1.86 B -1 1 0 pp512 2788.83 ± 4.85 2792.44 ± 5.55 +0.1%
llama 1B Q8_0 1.84 GiB 1.86 B -1 1 0 tg128 119.51 ± 0.52 121.10 ± 0.53 +1.3%
deepseek2 30B.A3B Q3_K - Small 12.37 GiB 29.94 B -1 1 0 pp512 65.16 ± 0.18 65.25 ± 0.28 +0.1%
deepseek2 30B.A3B Q3_K - Small 12.37 GiB 29.94 B -1 1 0 tg128 39.09 ± 0.15 39.36 ± 0.13 +0.7%

Nvidia RTX 3090:

model size params ngl fa mmap test t/s (before) t/s (after) diff
llama 8B Q4_0 4.33 GiB 8.03 B -1 1 0 pp512 4737.72 ± 21.83 4699.74 ± 27.75 -0.8%
llama 8B Q4_0 4.33 GiB 8.03 B -1 1 0 tg128 148.43 ± 0.86 148.03 ± 0.73 -0.3%
llama 1B Q8_0 1.84 GiB 1.86 B -1 1 0 pp512 12239.14 ± 211.22 12196.23 ± 94.26 -0.4%
llama 1B Q8_0 1.84 GiB 1.86 B -1 1 0 tg128 307.75 ± 1.86 310.28 ± 2.08 +0.8%
deepseek2 30B.A3B Q3_K - Small 12.37 GiB 29.94 B -1 1 0 pp512 2241.77 ± 9.81 2228.30 ± 14.53 -0.6%
deepseek2 30B.A3B Q3_K - Small 12.37 GiB 29.94 B -1 1 0 tg128 123.54 ± 0.37 123.49 ± 0.51 -0.0%

@0cc4m

0cc4m commented Jun 17, 2026

Copy link
Copy Markdown
Contributor Author

Also this prototype is likely not thread-safe, as that CI failure suggests.

@jeffbolznv

Copy link
Copy Markdown
Contributor

I did a quick perf test and I think it was all just noise. If we can't see a meaningful speedup I'd prefer not to do this, I think it has been kind of fragile for cuda graphs to cache the tensor properties. I thought at some point there was a plan to have some kind of first-class ggml data structure that corresponded to a reusable graph, which would probably make this simpler for backends.

@inforithmics

inforithmics commented Jun 18, 2026 •

Copy link
Copy Markdown

I tried to evaluate this patch against mtp but unfortunatly llama-bench does not support this as far as I know. So i used the same llama-cli command prompt to see if there is a difference on Linux

llama-bench device output:
ggml_vulkan: 0 = AMD Radeon 780M Graphics (RADV PHOENIX) (radv) | uma: 1 | fp16: dot2 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 1 | matrix cores: KHR_coopmat

./llama-cli -hf unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q5_K_XL --spec-type draft-mtp -p "write a hello world programm"

PR:
[ Prompt: 39.7 t/s | Generation: 26.3 t/s ]
Main:
[ Prompt: 43.2 t/s | Generation: 25.3 t/s ]

so there is around 4% speed increase in tg/s

@inforithmics

inforithmics commented Jun 20, 2026 •

Copy link
Copy Markdown

I ran speed bench and there is no difference so my earlier bench was some fluke

data

Pr:
summary (elapsed=2739.76s)
category samples avg_prompt_t/s avg_pred_t/s avg_latency accept_rate


coding 10 147.70 25.05 25.661s 0.7292
humanities 10 41.78 22.89 24.988s 0.6363
math 10 48.93 23.76 23.713s 0.6834
qa 10 51.22 23.89 20.995s 0.6714
rag 10 207.46 25.60 23.748s 0.7587
reasoning 10 34.36 22.93 22.488s 0.6371
stem 10 36.89 22.76 22.615s 0.6314
writing 10 187.84 24.09 25.750s 0.6963
multilingual 10 94.38 26.51 18.232s 0.7906
summarization 10 56.06 24.52 21.344s 0.7052
roleplay 10 144.51 22.75 44.441s 0.6240
overall 110 95.56 24.07 24.907s 0.6803

Main:
Summary (elapsed=2743.04s)
category samples avg_prompt_t/s avg_pred_t/s avg_latency accept_rate


coding 10 148.39 25.01 25.693s 0.7292
humanities 10 41.89 22.90 24.982s 0.6363
math 10 48.92 23.66 23.815s 0.6834
qa 10 51.13 23.91 20.981s 0.6714
rag 10 206.91 25.53 23.809s 0.7587
reasoning 10 34.53 22.84 22.579s 0.6371
stem 10 36.92 22.74 22.632s 0.6314
writing 10 188.07 24.12 25.710s 0.6963
multilingual 10 93.68 26.52 18.232s 0.7906
summarization 10 55.67 24.49 21.377s 0.7052
roleplay 10 142.96 22.72 44.493s 0.6240
overall 110 95.37 24.04 24.937s 0.6803

python tools/server/bench/speed-bench/speed_bench_compare.py --baseline main.json --speculative pr.json

Comparison: baseline=main.json speculative=pr.json

category base_avg_pred_t/s spec_avg_pred_t/s decode_speedup base_avg_latency spec_avg_latency latency_speedup accept_rate
coding 25.01 25.05 1.00x 25.693s 25.661s 1.00x 0.7292
humanities 22.90 22.89 1.00x 24.982s 24.988s 1.00x 0.6363
math 23.66 23.76 1.00x 23.815s 23.713s 1.00x 0.6834
qa 23.91 23.89 1.00x 20.981s 20.995s 1.00x 0.6714
rag 25.53 25.60 1.00x 23.809s 23.748s 1.00x 0.7587
reasoning 22.84 22.93 1.00x 22.579s 22.488s 1.00x 0.6371
stem 22.74 22.76 1.00x 22.632s 22.615s 1.00x 0.6314
writing 24.12 24.09 1.00x 25.710s 25.750s 1.00x 0.6963
multilingual 26.52 26.51 1.00x 18.232s 18.232s 1.00x 0.7906
summarization 24.49 24.52 1.00x 21.377s 21.344s 1.00x 0.7052
roleplay 22.72 22.75 1.00x 44.493s 44.441s 1.00x 0.6240
overall 24.04 24.07 1.00x 24.937s 24.907s 1.00x 0.6803

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning Vulkan Issues specific to the Vulkan backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants