Repository navigation
Conversation
|
Some benchmarks, all on AMD EPYC 7302 server: AMD Radeon Pro VII:
Intel A770:
Nvidia RTX 3090:
|
|
Also this prototype is likely not thread-safe, as that CI failure suggests. |
|
I did a quick perf test and I think it was all just noise. If we can't see a meaningful speedup I'd prefer not to do this, I think it has been kind of fragile for cuda graphs to cache the tensor properties. I thought at some point there was a plan to have some kind of first-class ggml data structure that corresponded to a reusable graph, which would probably make this simpler for backends. |
|
I tried to evaluate this patch against mtp but unfortunatly llama-bench does not support this as far as I know. So i used the same llama-cli command prompt to see if there is a difference on Linux llama-bench device output: ./llama-cli -hf unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q5_K_XL --spec-type draft-mtp -p "write a hello world programm" PR: so there is around 4% speed increase in tg/s |
|
I ran speed bench and there is no difference so my earlier bench was some fluke dataPr: coding 10 147.70 25.05 25.661s 0.7292 Main: coding 10 148.39 25.01 25.693s 0.7292 python tools/server/bench/speed-bench/speed_bench_compare.py --baseline main.json --speculative pr.json Comparison: baseline=main.json speculative=pr.json
|
Overview
This is more of a prototype to see if reusing graphs/command buffers is possible and what difference it makes. I cannot see much of a performance difference at all from this, which suggests we're already overlaying GPU and CPU work well-enough that not much is to gain from reducing CPU work.
However, my list of devices to test on is limited, and it may also be advantageous to reduce the CPU load in this way, by not rerecording identical graphs each token.
@jeffbolznv Let me know what you think, whether you think this could be worth adding or not (and not necessarily this specific implementation, I didn't think it through enough yet)
Requirements