Skip to content

Feature Request: make adaptive or changeable kvarn decode split-128 crossover threshold for more performance #183

Description

@masel

Prerequisites

  • I am running the latest code. Mention the version if possible as well.
  • I carefully followed the README.md.
  • I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
  • I reviewed the Discussions, and have a new and useful enhancement to share.

Feature Description

In fattn-mma-kvarn-decode.cuh the split-128 kernel for n_q == 1, D != 128 is currently not allowed below n_kv = 32768

if (n_q == 1 && ((D == 128 && n_kv < 4096) ||
        (D != 128 && n_kv < 32768))) {
    return;
}

On smaller cards, such as the 5060 TI, some TG performance is therefore left on the table, particularly when MTP is enabled. Making this threshold adaptive, or at least adjustable, could help with testing and tuning, especially on these smaller cards. Maybe it's even worth adjusting it down on 3090 with drafting enabled.

Motivation

I did some performance testing and noticed a visible step in TG performance at 32768 context. This was tracked down by Claude, Opus 5.5, to the split-128 kernel threshold mentioned above. An overridable threshold via an environment variable was added like this:

--- a/ggml/src/ggml-cuda/fattn-mma-kvarn-decode.cuh
+++ b/ggml/src/ggml-cuda/fattn-mma-kvarn-decode.cuh
@@ -5,6 +5,8 @@
 #include "fattn-mma-kvarn-load.cuh"
 #include "mma.cuh"
 
+#include <cstdlib>
+
 using namespace ggml_cuda_mma;
 
 static constexpr int GGML_CUDA_FATTN_KVARN_DECODE_THREADS = 256;
@@ -969,8 +971,14 @@ static void ggml_cuda_fattn_kvarn_decode_consider(
     if constexpr (SPLIT_TOKENS == 128) {
         // A D128 CTA can assign its eight warps to both halves of the record;
         // wider heads keep the measured deep-context crossover.
+        // Experiment: GGML_KVARN_SPLIT128_MIN_KV overrides the crossover for
+        // wider heads; unset keeps the default of 32768.
+        static const int split128_min_kv = [] {
+            const char * value = getenv("GGML_KVARN_SPLIT128_MIN_KV");
+            return value != nullptr ? atoi(value) : 32768;
+        }();
         if (n_q == 1 && ((D == 128 && n_kv < 4096) ||
-                (D != 128 && n_kv < 32768))) {
+                (D != 128 && n_kv < split128_min_kv))) {
             return;
         }
     }

(GGML_KVARN_SPLIT128_MIN_KV=0 allows split-128 at any depth)

and benched afterwards. This gave the following results with llama-bench (no MTP) on RTX 5060 Ti 16 GB, driver 616.92, Windows 11, CUDA build for sm_120, beellama.cpp v0.4.8 @ 49df8fa, Qwen3.8-27B-UD-Q3_K_XL, -ctk/-ctv kvarn4, tg128 at the given depth

Run the same command twice, once without the variable and once with GGML_KVARN_SPLIT128_MIN_KV=0:

llama-bench -m Qwen3.8-27B-UD-Q3_K_XL.gguf -ngl 99 -fa on -ctk kvarn4 -ctv kvarn4 \
  -p 0 -n 128 -d 2000,10000,22000,30000,34000 -r 3 -o csv
depth default t/s threshold 0 t/s change kvarn_geometry_split_64 (default)
2000 27.48 ± 0.04 27.03 ± 0.03 −1.6 % 208
10000 26.58 ± 0.13 26.51 ± 0.06 −0.2 % 48
22000 24.56 ± 0.02 25.06 ± 0.06 +2.1 % 208
30000 23.41 ± 0.04 23.94 ± 0.06 +2.3 % 48
34000 23.29 ± 0.06 23.31 ± 0.01 ±0 (same kernel) 16

With a threshold of 0, the value of kvarn_geometry_split_64 is 0 at every depth, confirming that the override changes the route. At 34,000, both runs use the same kernel. This row acts as a control and demonstrates that the difference is not due to noise. The reason for those jumping kvarn_geometry_split_64 values is unknown to me.

Finer sweep with llama-server (own benchmark tool)

I also measured with llama-server over /completion with agentic-coding-style prompts, cold cache per point, 512 generated tokens, 3 repeats, temperature 0.6. Below is the change in decode time per step (predicted_ms / (predicted_n − draft_n_accepted)) with threshold 0 against the default. A negative value means split-128 is faster. The spread within each point was 0.1–0.3 %.

ctx no MTP MTP (--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.0)
1000 +0.9 % +0.9 %
2000 +1.6 % +0.8 %
4000 +0.7 % ±0.0 %
6000 +1.0 % −0.6 %
8000 +1.0 % −1.1 %
10000 +0.3 % / ±0.0 % −1.6 %
14000 ±0.0 % −3.5 %
18000 −1.0 % −4.1 %
22000 −2.3 % −4.8 %
26000 −2.4 % −5.9 %
30000 −2.3 % −6.8 %
34000–42000 ±0.0 % −0.2 … 0 %

At 10000 the no-MTP value was measured in short sweep and long sweep, hence the two values.

  • Without MTP the crossover is at ~14k–16k. Below that, split-64 is up to ~1.6 % faster.
  • With MTP the crossover moves down to ~4k, and the gain grows to ~7 % just below 32768. The reason for this is unknown. My guess is that the MTP draft head's single-token decodes also hit the n_q == 1 path, but I haven't verified that.

Tried to reproduce this on a RTX PRO 6000 Blackwell System via modal to rule out architectural effects, but the system had ~5% measurement noise so effects of 1-2% would be hidden. The ~7 % MTP effect seen on the 5060 Ti did not show up. I'm happy to run further tests as needed.

Possible Implementation

The hypothesis is that this is related to the SM count, as these might have been filled earlier by the split-64 grid. So I suggest a threshold by the device's SM count e.g. for non-MTP maybe 32768 * nsm / 82, or at least an environment override for further testing and feedback.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions