You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 32af709
Browse filesBrowse the repository at this point in the historyBrowse files
qwen4exp: fix speculative decoding and add native MTP (NextN) support
Speculative decoding on Qwen3.8-Flash-Next previously ran ~6x SLOWER than
plain decode. This branch fixes seven issues: the recurrent rollback ring was
not enabled for this arch (every spec step serialized the full recurrent state
through host RAM), the ring was granted only to MTP/EAGLE3-style types, the
qwen4exp conv-state write ignored the ring banks, the PLE n-gram history could
not rewind losslessly, the rejection path never restored a draft context that
cannot roll back, prompt-cache reuse desynced such drafts, and
--spec-draft-model with draft-mtp loaded the main model instead of the given
sidecar file.
It also implements the model's native MTP head end to end: an --mtp converter
export (mtp.* tensors -> an mtp- sidecar GGUF) and the draft-mtp graph. The
combiner keeps the hyper-connection residual streams distinct, which is what
the head was trained on (draft acceptance 0.87 vs 0.47 when collapsed).
Measured on Strix Halo (Ryzen AI Max+ 395, Vulkan/RADV), UD-Q3_K_XL target,
greedy: decode 23.6 -> 49.1 t/s (2.08x) at shallow context and 14.0 -> 31.5
t/s (2.24x) at 7.3k tokens, acceptance 0.785 at --spec-draft-n-max 6.
Correctness verified with a greedy identity oracle (spec output equals plain
decode, modulo replay FP-reduction-order noise).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
return t == COMMON_SPECULATIVE_TYPE_DRAFT_MTP || t == COMMON_SPECULATIVE_TYPE_DRAFT_EAGLE3 || t == COMMON_SPECULATIVE_TYPE_DRAFT_DFLASH || t == COMMON_SPECULATIVE_TYPE_DRAFT_DSPARK;
393
+
return t != COMMON_SPECULATIVE_TYPE_NONE;
389
394
});
390
395
391
-
return needs_rs_seq ? draft.n_max : 0u;
396
+
// n_max + 1: a verify batch holds the previously sampled token plus up to n_max
397
+
// drafts, and the checkpoint+replay path (used when the draft context cannot roll
398
+
// back) rewinds the target across the whole batch, one deeper than the rejected
0 commit comments