Prerequisites
Feature Description
In fattn-mma-kvarn-decode.cuh the split-128 kernel for n_q == 1, D != 128 is currently not allowed below n_kv = 32768
if (n_q == 1 && ((D == 128 && n_kv < 4096) ||
(D != 128 && n_kv < 32768))) {
return;
}
On smaller cards, such as the 5060 TI, some TG performance is therefore left on the table, particularly when MTP is enabled. Making this threshold adaptive, or at least adjustable, could help with testing and tuning, especially on these smaller cards. Maybe it's even worth adjusting it down on 3090 with drafting enabled.
Motivation
I did some performance testing and noticed a visible step in TG performance at 32768 context. This was tracked down by Claude, Opus 5.5, to the split-128 kernel threshold mentioned above. An overridable threshold via an environment variable was added like this:
--- a/ggml/src/ggml-cuda/fattn-mma-kvarn-decode.cuh
+++ b/ggml/src/ggml-cuda/fattn-mma-kvarn-decode.cuh
@@ -5,6 +5,8 @@
#include "fattn-mma-kvarn-load.cuh"
#include "mma.cuh"
+#include <cstdlib>
+
using namespace ggml_cuda_mma;
static constexpr int GGML_CUDA_FATTN_KVARN_DECODE_THREADS = 256;
@@ -969,8 +971,14 @@ static void ggml_cuda_fattn_kvarn_decode_consider(
if constexpr (SPLIT_TOKENS == 128) {
// A D128 CTA can assign its eight warps to both halves of the record;
// wider heads keep the measured deep-context crossover.
+ // Experiment: GGML_KVARN_SPLIT128_MIN_KV overrides the crossover for
+ // wider heads; unset keeps the default of 32768.
+ static const int split128_min_kv = [] {
+ const char * value = getenv("GGML_KVARN_SPLIT128_MIN_KV");
+ return value != nullptr ? atoi(value) : 32768;
+ }();
if (n_q == 1 && ((D == 128 && n_kv < 4096) ||
- (D != 128 && n_kv < 32768))) {
+ (D != 128 && n_kv < split128_min_kv))) {
return;
}
}
(GGML_KVARN_SPLIT128_MIN_KV=0 allows split-128 at any depth)
and benched afterwards. This gave the following results with llama-bench (no MTP) on RTX 5060 Ti 16 GB, driver 616.92, Windows 11, CUDA build for sm_120, beellama.cpp v0.4.8 @ 49df8fa, Qwen3.8-27B-UD-Q3_K_XL, -ctk/-ctv kvarn4, tg128 at the given depth
Run the same command twice, once without the variable and once with GGML_KVARN_SPLIT128_MIN_KV=0:
llama-bench -m Qwen3.8-27B-UD-Q3_K_XL.gguf -ngl 99 -fa on -ctk kvarn4 -ctv kvarn4 \
-p 0 -n 128 -d 2000,10000,22000,30000,34000 -r 3 -o csv
| depth |
default t/s |
threshold 0 t/s |
change |
kvarn_geometry_split_64 (default) |
| 2000 |
27.48 ± 0.04 |
27.03 ± 0.03 |
−1.6 % |
208 |
| 10000 |
26.58 ± 0.13 |
26.51 ± 0.06 |
−0.2 % |
48 |
| 22000 |
24.56 ± 0.02 |
25.06 ± 0.06 |
+2.1 % |
208 |
| 30000 |
23.41 ± 0.04 |
23.94 ± 0.06 |
+2.3 % |
48 |
| 34000 |
23.29 ± 0.06 |
23.31 ± 0.01 |
±0 (same kernel) |
16 |
With a threshold of 0, the value of kvarn_geometry_split_64 is 0 at every depth, confirming that the override changes the route. At 34,000, both runs use the same kernel. This row acts as a control and demonstrates that the difference is not due to noise. The reason for those jumping kvarn_geometry_split_64 values is unknown to me.
Finer sweep with llama-server (own benchmark tool)
I also measured with llama-server over /completion with agentic-coding-style prompts, cold cache per point, 512 generated tokens, 3 repeats, temperature 0.6. Below is the change in decode time per step (predicted_ms / (predicted_n − draft_n_accepted)) with threshold 0 against the default. A negative value means split-128 is faster. The spread within each point was 0.1–0.3 %.
| ctx |
no MTP |
MTP (--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.0) |
| 1000 |
+0.9 % |
+0.9 % |
| 2000 |
+1.6 % |
+0.8 % |
| 4000 |
+0.7 % |
±0.0 % |
| 6000 |
+1.0 % |
−0.6 % |
| 8000 |
+1.0 % |
−1.1 % |
| 10000 |
+0.3 % / ±0.0 % |
−1.6 % |
| 14000 |
±0.0 % |
−3.5 % |
| 18000 |
−1.0 % |
−4.1 % |
| 22000 |
−2.3 % |
−4.8 % |
| 26000 |
−2.4 % |
−5.9 % |
| 30000 |
−2.3 % |
−6.8 % |
| 34000–42000 |
±0.0 % |
−0.2 … 0 % |
At 10000 the no-MTP value was measured in short sweep and long sweep, hence the two values.
- Without MTP the crossover is at ~14k–16k. Below that, split-64 is up to ~1.6 % faster.
- With MTP the crossover moves down to ~4k, and the gain grows to ~7 % just below 32768. The reason for this is unknown. My guess is that the MTP draft head's single-token decodes also hit the
n_q == 1 path, but I haven't verified that.
Tried to reproduce this on a RTX PRO 6000 Blackwell System via modal to rule out architectural effects, but the system had ~5% measurement noise so effects of 1-2% would be hidden. The ~7 % MTP effect seen on the 5060 Ti did not show up. I'm happy to run further tests as needed.
Possible Implementation
The hypothesis is that this is related to the SM count, as these might have been filled earlier by the split-64 grid. So I suggest a threshold by the device's SM count e.g. for non-MTP maybe 32768 * nsm / 82, or at least an environment override for further testing and feedback.
Prerequisites
Feature Description
In fattn-mma-kvarn-decode.cuh the split-128 kernel for
n_q == 1,D != 128is currently not allowed belown_kv = 32768On smaller cards, such as the 5060 TI, some TG performance is therefore left on the table, particularly when MTP is enabled. Making this threshold adaptive, or at least adjustable, could help with testing and tuning, especially on these smaller cards. Maybe it's even worth adjusting it down on 3090 with drafting enabled.
Motivation
I did some performance testing and noticed a visible step in TG performance at 32768 context. This was tracked down by Claude, Opus 5.5, to the split-128 kernel threshold mentioned above. An overridable threshold via an environment variable was added like this:
(
GGML_KVARN_SPLIT128_MIN_KV=0allows split-128 at any depth)and benched afterwards. This gave the following results with llama-bench (no MTP) on RTX 5060 Ti 16 GB, driver 616.92, Windows 11, CUDA build for sm_120, beellama.cpp v0.4.8 @ 49df8fa, Qwen3.8-27B-UD-Q3_K_XL, -ctk/-ctv kvarn4, tg128 at the given depth
Run the same command twice, once without the variable and once with
GGML_KVARN_SPLIT128_MIN_KV=0:kvarn_geometry_split_64(default)With a threshold of 0, the value of kvarn_geometry_split_64 is 0 at every depth, confirming that the override changes the route. At 34,000, both runs use the same kernel. This row acts as a control and demonstrates that the difference is not due to noise. The reason for those jumping kvarn_geometry_split_64 values is unknown to me.
Finer sweep with llama-server (own benchmark tool)
I also measured with llama-server over
/completionwith agentic-coding-style prompts, cold cache per point, 512 generated tokens, 3 repeats, temperature 0.6. Below is the change in decode time per step (predicted_ms / (predicted_n − draft_n_accepted)) with threshold 0 against the default. A negative value means split-128 is faster. The spread within each point was 0.1–0.3 %.--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.0)At 10000 the no-MTP value was measured in short sweep and long sweep, hence the two values.
n_q == 1path, but I haven't verified that.Tried to reproduce this on a RTX PRO 6000 Blackwell System via modal to rule out architectural effects, but the system had ~5% measurement noise so effects of 1-2% would be hidden. The ~7 % MTP effect seen on the 5060 Ti did not show up. I'm happy to run further tests as needed.
Possible Implementation
The hypothesis is that this is related to the SM count, as these might have been filled earlier by the split-64 grid. So I suggest a threshold by the device's SM count e.g. for non-MTP maybe 32768 * nsm / 82, or at least an environment override for further testing and feedback.