Repository navigation
qwen4exp : add NextN/MTP draft head (--spec-type draft-mtp) for Qwen3.8-Flash-Next - #27836
rmonsurate wants to merge 3 commits into
Conversation
Adds the MTP head's own hyper-connection mixer tensor names and lists the NextN tensors under the qwen4exp architecture.
Adds --spec-type draft-mtp support for Qwen3.8-Flash-Next. The MTP head folds the next token's embedding into the trunk's wide hyper-connection residual, runs one trunk-style block (dense attention + MoE) over it, and collapses the result with its own mixer before reusing the trunk's LM head. - read nextn_predict_layers so n_layer() excludes the MTP block - load the trailing block through the existing trunk path: is_recr() and is_ple() are already false past the trunk, so it needs no special casing - eh_proj fuses the checkpoint's fc_embedding and fc_hidden side by side, so one matmul computes fc_embedding@e + fc_hidden@h - the head carries its own hyper-connection mixer, mirroring the trunk's hc_head_*, which stands in for the output norm qwen4exp does not have - export the wide pre-collapse residual as t_h_nextn from both graphs, so the driver can feed it back for the next draft step - route MTP contexts to a plain KV cache filtered to the trailing layer The draft block attends densely for now: the trunk's QSA only prunes context past a 2048-token budget, so dense is a numerical superset and drafts are verified either way. Indexer tensors are still loaded.
The MTP block is one trunk-shaped block (dense attention + MoE wrapped in hyper-connections) plus a head-level combiner, so once _QwenMtpMixin renames mtp.layers.0.* to the trailing block index its tensors ride the existing qwen4exp mappings unchanged. Two head-level pieces need handling: - fc_embedding and fc_hidden fuse into the eh_proj the shared NextN code expects, since W_e@e + W_h@h == [W_e|W_h] @ concat(e, h) - mtp.hyper_connection_mixer.* is the head's own copy of the trunk's hc_head_* output mixer, unindexed in the checkpoint and per-block in the GGUF compress_ratios is read with length block_count, so it gains a trailing 0 for the MTP block, which attends densely. --no-nextn drops the head; --mtp exports it on its own.
|
Hi @rmonsurate, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
AMD HIP validation + tuning report (2× Radeon AI PRO R9700, gfx1201, TheRock 7.14)Thanks for this PR — it fixes the exact decode crash we hit on the #27742 merge (the MTP works on AMD HIP. Load ( Notes for the converter / file formatThe only community export we could find ( Benchmarks (Q4_K_XL 5-shard, q8 KV, 16K ctx, temp 0, thinking off, 3 reps, harness-measured)
Happy to share the full per-rep matrix or the graft script if useful. Thanks again for the fix. |
|
I could not get this to work, so i asked qwen 3.8 to help me a bit. Now it is working for me, see https://github.com/freqmod/llama.cpp/tree/qwen4exp-mtp if you want to be inspired. i got it to patch the source code and the script to build the files so i don't have to download the whole model to convert it NB: this is posted here mostly if anyone tinkering wants to see and use it while the pr gets approved. It is not following the llama.cpp ai policy as it is mostly done by qwen 3.8 27t so it is probably not useful for merging into the review. I quickly skimmed trough the diff and it looks reasonable to me. also it works on my computer whatever that is worth. |
|
Can't get this to work either, it seems like the GGUFs do not come with MTP heads enabled or something, anybody knows which GGUF I should download to get this working ? |
Vulkan validation report (gfx1151 / Strix Halo, RADV) — works, but draft path is a net loss on this backendBuilt What works:
What doesn't (Vulkan-specific):
Question: is the draft context expected to checkpoint/reuse the recurrent state across draft steps on rejection, or is the full replay by design for hybrid archs? Happy to run additional Vulkan datapoints if useful. |
|
Would the work this guy is doing be of any interest https://github.com/Nathanw1014/llama.cpp/commits/strix-halo-vulkan It just works out of the box, I'm getting 21-30t/s with MTP on a EVO X2 Windows 11 128GB unified memory with Vulkan. Quite impressive as I wasn't expecting to get that performance at all. I am using unsloth/Qwen3.8-Flash-Next-UD-Q4_K_XL with a separate MTP model Qwen3.8-Flash-Next-MTP-Q4_K_M.gguf I tried using this PR but couldn't even load the MTP draft |
|
This doesn't work for me, but the reason is trivial: it looks for blk.0, whereas all GGUFs on hf (which correctly follow convention) ship blk.38. This patch fixes it (cherry-pick this commit): crusaderky@a82a58a (beware: unreviewed AI slop). This branch + the above commit work for me with on CUDA RTX3090, using this presets.ini: [Qwen3.8-Flash-Next]
hf = unsloth/Qwen3.8-Flash-Next-GGUF:UD-IQ4_XS
ngl = 99
n-cpu-moe = 48
load-mode = mmap
jinja = true
ctx-size = 262144
flash-attn = on
kv-unified = true
cache-type-k = q8_0
cache-type-v = q8_0
# MTP drafter
spec-draft-hf = agentionai/Qwen3.8-Flash-Next-MTP-Q8_0-GGUF:Q8_0
spec-type = draft-mtp
spec-draft-ngl = 99
spec-draft-n-max = 3
spec-draft-p-min = 0.6
# Thinking mode
temperature = 1.0
top-p = 0.95
top-k = 20
min-p = 0.0
presence-penalty = 0.0
repeat-penalty = 1.0 |
|
Data point supporting your combiner warning, from Strix Halo (Ryzen AI Max+ 395 / gfx1151, ROCm 7.1): We tested an independent MTP port for qwen4exp (
Identical at Since that branch's issues are disabled, leaving the repro here as it seems relevant to "the combiner must be run per hyper-connection stream... If you do mean pooling first, the acceptance rate drops catastrophically." Happy to run this PR on gfx1151 (128 GB Strix Halo, UD-IQ4_XS) and report accept rates / tok/s / long-prompt correctness if useful — that hardware currently has no published numbers. |
|
Follow-up from the gfx1151 side: I tried to produce and run a standalone draft sidecar with this PR's Setup: fetched only the MTP-relevant tensors from Issue 1 — keep = name in (
"model.embed_tokens.weight", "model.norm.weight", "lm_head.weight",
"embed_tokens.weight", "norm.weight",
) or name.startswith("model.hyper_connection_mixer.")(The checkpoint names it Issue 2 — the exported sidecar can't be loaded standalone. The export writes
Question: what is the intended usage — was your M3 Max test run with the draft head still merged in the target gguf (self-speculative, no Happy to test any layout on Strix Halo — target UD-IQ4_XS, 128 GB, ROCm 7.1. The ngram-mod baseline here is 33 tok/s on file-edit workloads, so an MTP acceptance rate like your 89% would be a big deal on this hardware. |
unsloth ships this checkpoint at 83.8 gib (ud-q3_k_xl) / 87.2 gib (ud-iq4_xs) vs our 115.5 gib iq4_xs from the same f16. the gap is the 51b n-gram embedding table, which llama-quantize keeps at high precision. that, not the engine or the context size, is what made mtp not fit in 128g. correction posted to the engramhalo thread. also records that qwen4exp mtp now exists upstream (ggml-org/llama.cpp#27836, apepojken/llama.cpp@32af70900) and that our sidecar renumbering used the wrong contract.
|
Delivering the promised gfx1151/ROCm numbers — this PR + @crusaderky's loader fix works on Strix Halo, and it composes with ngram-mod and with the hipCUB TOP_K path. Setup: Ryzen AI Max+ 395 / Radeon 8060S, ROCm 7.1, target Correctness first: greedy output at 2.7k-token prompts is clean in every configuration (this is where the earlier independent MTP port degenerated into multilingual noise), and arithmetic/tool-calling/vision all check out at 24k-token context. Decode tok/s on real coding workloads (llama-server, greedy, 718-line Python file; A = rewrite file with small edit, B = targeted bugfix, C = write new code, D = prose):
Baseline without any speculation is 16.8 tok/s, so the full stack is 2.8x on file rewrites at 8k and +58% on code-writing at 24k. (hipCUB runs with B's variance across runs (24.7–40.2) tracks generation-content differences between builds at temp 0, not a regression — flagging it so nobody reads the last column pair as a CUB penalty. This configuration is now our production stack. Thanks @rmonsurate and @crusaderky — happy to run any follow-up variant on this hardware. |
|
Follow-up to my Vulkan report above — we've now run the same grafted model through HIP/ROCm 7.2.3 on the same gfx1151 box (containerized, MTP works great on ROCm (same spot where Vulkan collapsed):
Two operational notes for anyone reproducing on Strix Halo:
One honest caveat: in our end-to-end agentic workload (browser automation + tool calls, 97-99% prompt cache hits), total run time was statistically identical between ROCm+MTP and Vulkan no-spec at equal power limits — the LLM-side gains dilute into tool/navigation time. The +17-39% is real for LLM-bound workloads (batch, long prefills, sustained generation), which is where this PR shines. Thanks again for the head — acceptance numbers held up beautifully across two backends. |
|
For anyone wanting to run this without assembling the pieces by hand, the working combination is now published:
@gabrielfreire this combination is what got the draft loading on our end — the load failure you hit is the detached-head issue crusaderky's commit fixes. |
|
Follow-up datapoint from the Windows/CUDA side (RTX 3090 24 GB, 96 GB DDR4, The combiner fix is confirmed on CUDA. Draft acceptance: 0.84–0.91 on code, 0.75–0.82 on step-by-step reasoning, 0.62 on German prose — versus 0.33–0.40 we measured earlier on a mean-pooling-combiner fork (same GGUF, same prompts). Keeping the hyper-connection streams distinct is clearly the difference. But on a 24 GB card the speedup mostly evaporates: with It composes beautifully with timadinorth's MoE expert-residency split (timadinorth#1 — hot experts in VRAM instead of whole layers): with ~74% of routed traffic served from VRAM, the verify batches get cheap and the same box reaches 24.1–26.4 t/s on code (+38–52% over baseline) at 0.86–0.91 acceptance. Combined branch for reproduction: https://github.com/mjungnickel18/llama.cpp/tree/qwen4exp-mtp-plus-moe-residency (MTP branch + residency commit cherry-picked, flags in the branch's service script: Happy to re-test on this setup once the head lands here upstream. |
|
Adding some data for Apple Silicon. M5 Max 128GB, Unsloth UD-IQ4_XS with the standalone MTP head GGUF merged into the target's file set (nextn tensors as a 4th split, since current converts drop them). Prompts are C++ source-code continuation; prose lands a few t/s lower with the same ordering.
Short context: 41.6 t/s no spec → ~70 with One finding: combining with draftless ngram ( |
|
ON_DEVICE change is on my fork: JayToltTech#1. Six lines in server-context.cpp on current master. Three backends now. Vulkan gfx1151: 4.33 → 16.08 tok/s at 70k. @mjungnickel18 CUDA: +61%. @ovidiu-morar Metal: neutral, flag verified not a no-op. Win where per-round host round-trips are expensive, neutral where they aren't, no regression anywhere tested. Not submitting it upstream myself. Anyone who wants it, take it as your own — no credit needed. |
|
Cherry-picked the three commits here onto current master ( MTP now works on current master, but it needed one fix that lives outside this PR — flagging it here since the rebased version will likely hit the same thing. Root cause of the
|
Qwen3.8-Flash-Next ships an MTP block in the checkpoint that the converter was dropping, so the model had no speculative path at all. Three pieces: - ggml-org#27836: the qwen4exp NextN/MTP draft head. Converter export (fc_embedding|fc_hidden fuse into the single eh_proj the shared NextN code expects), the nextn.hc_head_* tensors that stand in for the output norm qwen4exp does not have, and the LLM_GRAPH_TYPE_DECODER_MTP graph. Resolved against the fork's per-layer n_ff_exp accessor. - ggml-org#27210: adaptive draft depth, --spec-type draft-mtp-adaptive. Carries a delta-net fix that matters well beyond the adaptive path: build_conv_state was emitting a snapshot slot for every one of the n_rs_seq + 1 rollback depths, including the ones no rollback inside the batch can reach. Decode is one token, so all but one slot repeated the pre-batch state; the bound turns n_rs_seq into free headroom instead of a per-layer kernel-launch tax. - [fork] mtp_only/trunk_only probing in qwen4exp's load_arch_tensors, following the bailingmoe3/deepseek2 pattern. The draft head can now ship as its own GGUF: quantized apart from the trunk and pinned to the head GPU with --model-draft + --device-draft, rather than riding the trunk's tensor-split out onto the RPC fabric. Trunk-only tensors (hc_head_*, the PLE table, blk.0..n-1) become NOT_REQUIRED when the file has no blk.0, and the nextn block likewise when the file has no eh_proj. token_embd and output stay required in both halves, since qwen4exp sets mtp_use_dedicated_embeddings=false and the draft graph reuses them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U3H5motr51eTWujztSXykc
|
Windows/CUDA crash with integrated Qwen3.8 MTP head: missing F32/F16 binary-op support I tested this PR at: • PR head: 1d8de7c The model itself loads correctly and the MTP draft context is created, but speculative initialization crashes because the MTP graph mixes F32 activations with F16 norm/mixer tensors. CPU reproduction When forcing the integrated MTP block (blk.48.*) to CPU: This happens repeatedly during MTP graph initialization. ggml-cpu/binary-ops.cpp has handlers for: but not: and also no: Adding those cases fixes the CPU-side failure. CUDA reproduction With the MTP block resident on CUDA, initialization crashes with: The relevant dispatcher currently does: if (src0->type == GGML_TYPE_F32 && dst->type == GGML_TYPE_F32) {
op()(src0, src1, dst,
(const float *) src0_dd,
(const float *) src1_dd,
(float *) dst_dd,
stream);
}This does not check src1->type. If src1 is F16, it is still cast to float *. Its stride can therefore be valid for 2-byte F16 elements but invalid for 4-byte float, causing: GGML_ASSERT(nb10 % sizeof(src1_t) == 0)Adding explicit mixed-type dispatch fixes it: if (src0->type == GGML_TYPE_F32 &&
src1->type == GGML_TYPE_F32 &&
dst->type == GGML_TYPE_F32) {
op()(src0, src1, dst,
(const float *) src0_dd,
(const float *) src1_dd,
(float *) dst_dd,
stream);
} else if (src0->type == GGML_TYPE_F32 &&
src1->type == GGML_TYPE_F16 &&
dst->type == GGML_TYPE_F32) {
op()(src0, src1, dst,
(const float *) src0_dd,
(const half *) src1_dd,
(float *) dst_dd,
stream);
}I also added the equivalent support for: For the fused CUDA bin-broadcast path, the corresponding handlers are: launch_bin_bcast_pack<op, float, half, float>(...)and: launch_bin_bcast_pack<op, half, half, float>(...)qwen4exp graph I also made the relevant MTP norm scaling type-safe in src/models/qwen4exp.cpp. For the shared hyper-connection mixer: xn = ggml_mul(
ctx0,
xn,
w_norm->type == xn->type
? w_norm
: ggml_cast(ctx0, w_norm, xn->type)
);For the MTP hidden-state norm: h_norm = ggml_mul(
ctx0,
h_norm,
ggml_cast(ctx0, layer.nextn.hnorm, h_norm->type)
);Result After rebuilding with these changes, the exact same configuration now successfully initializes the integrated GPU MTP context. Before the fix: After the fix: So this was reproducibly the difference between crash and successful startup on CUDA/Blackwell. The working model configuration reports: and the integrated MTP layer is successfully kept on CUDA. I think there are two possible fixes:
The generic backend fix seems useful beyond this specific model, since the current CUDA path can silently reinterpret an F16 src1 as float * whenever src0 and dst are F32. I can provide the exact local diff if useful. |
Ports upstream PR ggml-org#27836 (open, unmerged) on top of this fork's qwen4exp support: the MTP head folds the next token's embedding into the trunk's wide hyper-connection residual, runs one trunk-shaped block (dense attention + MoE) over it, and collapses the result with its own hyper-connection mixer before reusing the trunk's LM head. Also updates the converter and gguf-py tensor mappings so conversion/qwen4exp.py --mtp can export the head. Also adds mtp_only/trunk_flags handling to load_arch_tensors, which the ported PR didn't include: without it, loading a standalone MTP-only checkpoint (all trunk tensors absent) fails outright instead of tolerating their absence, unlike the equivalent qwen35.cpp path. Detects MTP-only via the absence of blk.0.hc_attn_norm.weight, a tensor every trunk layer carries. Guards the embedding and LM-head fallbacks with a clear assert instead of a null-pointer crash when a checkpoint has neither a dedicated NextN embedding/head nor the trunk's own. Verified: full build, test-llama-archs shows no regressions. Loading a real Qwen3.8-Flash-Next MTP draft checkpoint (mtp-Qwen3.8-Flash- Next-shared-Q4_K_M.gguf) now gets past tensor loading correctly (the mtp_only path works) and fails with the new clear assertion rather than "token_embd.weight not found": that specific checkpoint has neither nextn.embed_tokens nor a trunk token_embd.weight anywhere, confirmed by inspecting its raw safetensors source directly (15 tensors total, no embedding table in any form). That's a model packaging gap in that specific file, not a code gap.
…ive test Adds draft_filename/spec_type=draft-mtp/spec_draft_n_max to the test profile, to actually benchmark Unsloth's advertised 1.3-1.7x MTP speedup against the 25 tok/s baseline measured earlier this session. Requires a custom-built llama-server (rmonsurate/llama.cpp@qwen4exp-mtp, commit 1d8de7c1 -- ggml-org/llama.cpp#27836, still a draft PR, not in any official ghcr.io/ggml-org/llama.cpp tag) built and confirmed live on circe this session: --help on the custom binary advertises --spec-type draft-mtp, which stock server-cuda-b10666 does not. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Live-tested on circe 2026-09-04: neither MTP companion file (shared or non-shared) actually loads on rmonsurate/llama.cpp@qwen4exp-mtp commit 1d8de7c1. Both crash the server fatally with a check_tensor_dims error on a different missing tensor each time. Commented out with the full failure detail rather than deleted, so a future attempt (e.g. against a newer commit once ggml-org/llama.cpp#27836 moves past draft) starts from what's already known instead of re-discovering it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
qwen4exp: NextN/MTP draft head (port of upstream PR ggml-org#27836) + follow-up fixes (ggml-org#27941)
The external "shared MTP module" draft file for Qwen3.8-Flash-Next
(Unsloth's mtp-*.gguf, and POM's own EngineConfig.mtp_draft_path path)
never attached: load_arch_tensors required the full trunk and the
model-level hc_head_{norm,down,up} unconditionally, so a draft-only
file (last block + nextn.* + token_embd/output, no trunk) failed at
the very first model-level tensor with check_tensor_dims: tensor
'output_hc_norm.weight' not found. Even past that, the architecture
never declared layer.nextn.* tensors or a DECODER_MTP graph class at
all - qwen4exp.cpp only exported h_nextn from the target side
(1b9c07a), never built the draft head's own graph.
Ports the mtp_only pattern already proven by qwen35.cpp/qwen3next.cpp/
deepseek4.cpp: load_arch_tensors now detects a nextn-only file (probes
blk.0.hc_attn_norm.weight) and marks the trunk and the model-level
hc_head_* TENSOR_NOT_REQUIRED in that case; the tensor loop now walks
the tail (n_layer..n_layer_all) too, declaring layer.nextn.{enorm,
hnorm,eh_proj,hc_head_norm,hc_head_down,hc_head_up,embed_tokens,
shared_head_head}; build_arch_graph dispatches to a new graph_mtp
class for LLM_GRAPH_TYPE_DECODER_MTP, adapted from the real upstream
implementation (ggml-org#27836) to this fork's TENSOR_ALLOW_
RESHAPE storage convention for the hc_*_norm gammas. Also fixes the
main graph's h_nextn export to the wide pre-mixer hyper-connection
residual (n_embd*hc) instead of the already-collapsed n_embd value -
exporting the narrow form would have silently starved the draft head
of most of its input instead of crashing, per upstream's own note
("mean pooling first drops acceptance catastrophically").
mtp_on_hybrid_qwen (llama-model.cpp) now includes LLM_ARCH_QWEN4EXP
so an MTP-context memory alloc on this hybrid architecture gets the
plain per-tail-layer KV cache instead of the full hybrid one.
Verified with a synthetic minimal qwen4exp GGUF matching the real
draft's shape (last block + nextn.* + token_embd/output, no trunk, no
model-level hc_head_*): reproduced the exact production failure before
this change, and confirmed llama_init_from_model(ctx_type=MTP) now
reserves the graph_mtp graph successfully after it (crates/llama-engine
tests, rdma-p2p repo). Not verified: the real Unsloth draft file, GPU
backends (only compiled/tested CPU-only here), and the actual ksia
deploy.
Co-authored-by: Claude <noreply@anthropic.com>
Upstream absorbed the fork's two largest features in this window: - c061df1 (ggml-org#29761) adds Qwen4Exp MTP. Structurally the same design the fork carried (both descend from ggml-org#27836): mtp_only/trunk_flags, graph_mtp, the NEXTN_HC_HEAD_* tensors and the hc-wide residual through eh_proj. Upstream's version is taken whole; only the shared-sidecar support is re-applied on top ([TAG_QWEN4EXP_SHARED_MTP], 34 lines: tok_embd/output optional when the trunk is absent, borrowed from cparams.ctx_other in graph_mtp). - 649dcb1 (ggml-org#27773) adds GLM-5.3-Flash as arch 'glm5-next' with its own src/models/glm5-next.cpp and vision projector 'glm5v'. The fork's 'glm5next' (from draft ggml-org#27752) and projector 'glm5next' are kept beside it: existing GGUFs carry the old arch string. Both arch cases and both projector cases stand side by side. Dropped: the multi-token QSA gather ([TAG_QSA_GATHER_VERIFY]) and its two fixtures. Upstream rewrote the QSA path (889edf4 halves the indexer score memory, 4e2713c rebuilds the masks, build_attn_qsa has a new signature), so the gather no longer applies textually and has to be re-derived. Two upstream bugs fixed on the way, both only visible on a virtual model: - qwen4exp's mtp_only probe used get_weight(), which always answers nullptr without files, so every virtual qwen4exp marked its whole trunk NOT_REQUIRED and aborted at the first tensor's buft. Gated on !ml.files.empty(). - create_tensor_qkv's skip constant carried TENSOR_SKIP_IF_VIRTUAL only on the fused QKV, not on the separate Q/K/V, so a skipped MTP block aborted the same way. Added to the constant. Also ported: the fork's [TAG_NAN_CHECK]-adjacent test helpers to common_batch / llama_process, llm_graph_input_mem_hybrid_kpool::can_reuse to upstream's llm_graph_input_rs (s_copy_extra is gone), and the hipCUB argsort guard around upstream's new shared-memory bitonic bound. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Overview
Adds the NextN/MTP draft head for qwen4exp. Tested in the Fabric harness using Qwen3.8-Flash-Next. This is a follow-up PR to #27742.
The combiner must be run per hyper-connection stream on the wide hidden state. If you do mean pooling first, the acceptance rate drops catastrophically — other people might hit this, so just leaving this here.
Converter support is included (
--mtp).Additional information
Measured on M3 Max 128GB, UD-IQ4_XS (temp-0 output byte-identical with MTP on/off):
--spec-draft-n-max 2--spec-draft-n-max 3Prior art: #27739.
Requirements