Skip to content

spec : add DFlash2 support (cherry-pick of ggml-org/llama.cpp#27816) - #261

Merged
bri-prism merged 1 commit into
PrismML-Eng:prismfrom
NakliTechie:dflash2-on-prism
Sep 25, 2026
Merged

bri-prism merged 1 commit into
PrismML-Eng:prismfrom
NakliTechie:dflash2-on-prism

Conversation

@NakliTechie

Copy link
Copy Markdown

Overview

Cherry-pick of upstream ggml-org#27816 (DFlash2: local convolution + candidate selector, merge commit b10f9ca) onto prism. Authorship is kept; the only local edits are formatting in src/llama-arch.cpp to match this branch.

With it, --spec-type draft-dflash loads DFlash2 drafters (the z-lab/Qwen3.8-27B-DFlash2 family) against Ternary Bonsai 2 27B. The borrowed tok_embd / output Hadamard handling from #210 covers the DFlash2 path; no extra change was needed.

Additional information

Tested on NVIDIA L4 (sm_89, CUDA 12.8), llama-server, Ternary-Bonsai-2-27B GGUF target + a DFlash2 Q4_K_M drafter, greedy, one slot:

prompt tokens tok/s draft accepted
email (prose) 251 31.2 149/714 (21%)
code (Python class + tests) 512 65.6 410/703 (58%)
story (prose) 512 28.9 282/1594 (18%)

Target PQ2_0, --spec-draft-n-max 7. Plain decode of the same target on the same GPU is about 30 tok/s, so code runs about 2x and prose about 1x. A larger run (HumanEval, MBPP, GSM8K, MT-Bench, thinking on/off, pass@1 plain vs speculative) is in progress; I will post it here.

The drafter used is a re-fit of z-lab's drafter to the ternary target (naklitechie/Qwen3.8-27B-DFlash2-ternary-bonsai2). Greedy output with speculation is not byte-identical to plain decode: splits happen only where the plain run's top-2 logprobs are within 0.03 nats (batched verify vs single-row kernels).

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. The upstream commit is unchanged apart from rebasing. Claude Code (Anthropic) did the cherry-pick, the CUDA build and the benchmark runs.

…gml-org#27342) (ggml-org#27816)

* spec : add DFlash2 support (local convolution + candidate selector) (ggml-org#27342)

* support DFlash2

* Add p_min in DFlash2

Assisted-by: Claude Opus 5

* Revert unnecessary changes

Assisted-by: Claude Opus 5

* Revert draft sampling in rejection sampling

Assisted-by: Claude Opus 5

* Refactor code structure

Assisted-by: Claude Opus 5

* Delete embedding scaling

Assisted-by: Claude Opus 5

* Gate output transforms on DFlash2

Assisted-by: Claude Opus 5

* Optimize Dflash 2 cost

Assisted-by: Claude Opus 5

* Avoid using atoi

Assisted-by: Claude Opus 5

* Modify comments

Assisted-by: Claude Opus 5

* Move llama_model_dflash_selector_top_k to llama-ext.h

Assisted-by: Claude Opus 5

* Formatting

Assisted-by: Claude Opus 5

* Apply patch to fix the mrope bug

Assisted-by: Claude Opus 5

* fix ci

Assisted-by: Claude Opus 5

* Fix graph number calculation

Assisted-by: Claude Opus 5

* rename hid and unary

Assisted-by: Claude Opus 5

---------

Co-authored-by: Jian Chen <jianchen0311@gmail.com>
Co-authored-by: Xuan-Son Nguyen <son@huggingface.co>

* revert top-k.cu changes

---------

Co-authored-by: Zihan Zhang <tiancaizhangdaxian@sjtu.edu.cn>
Co-authored-by: Jian Chen <jianchen0311@gmail.com>
(cherry picked from commit b10f9ca)
@NakliTechie

Copy link
Copy Markdown
Author

Benchmark on this branch, as promised in the description.

NVIDIA L4 (sm_89, CUDA 12.8), llama-server, greedy, batch 1, target Ternary-Bonsai-2-27B-PQ2_0.gguf, DFlash2 drafter Q4_K_M, --spec-draft-n-max 7. Speedup is decode tok/s vs the fastest plain config on the same GPU (PTQ1_0, no speculation).

thinking off n plain tok/s DFlash2 tok/s speedup pass@1 / EM, PQ2_0 plain -> DFlash2
GSM8K 100 31.4 67.6 2.15x 0.94 -> 0.93
MBPP (sanitized) 100 31.6 68.4 2.16x 0.80 -> 0.80
MATH-500 100 30.6 67.8 2.22x 0.76 -> 0.75
MT-Bench turn 1 80 31.3 42.9 1.37x -
  • Thinking on: 1.25-1.63x (40-prompt subsets).
  • --spec-type ngram-mod on the same sets: 0.92-0.95x.
  • A repeat on a second L4 machine accepted the same drafts (acceptance and tau identical to 3 decimals); speed within 3%.
  • Greedy output with speculation is not byte-identical to plain decode; splits happen at near-tied tokens, and accuracy differs by at most one problem per set.

Per-sample outputs, summaries and scripts: https://huggingface.co/datasets/naklitechie/bonsai2-dflash2-bench

@bri-prism bri-prism left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Checked this against upstream and ran it end to end on Metal and two RTX cards. Looks good to me.

Cherry-pick. All 15 files match b10f9ca hunk for hunk. The only differences are two dropped blank lines and one untouched LLM_KV_SHORTCONV_L_CACHE line in src/llama-arch.cpp that got reformatted. The DFlash2 logits go through build_lora_mm(output, ...), and the new get_rows calls only read the drafter's own selector tables, so the target's Hadamard transforms from #210 apply. #210 looks the transforms up by tensor pointer, so the tied output from #257 is covered too.

Tests. The new TOP_K cases pass 517/517 on Metal (M5 Pro), an RTX 4090 and an RTX 5090.

End to end. Ternary Bonsai 2 27B target, llama-server, one slot, greedy, --spec-draft-n-max 7, 256 tokens. Three DFlash2 drafters: yours (r3 Q4_K_M), ProCreations/Ternary-Bonsai-2-27B-DFlash2 (Q8_0) and incoai/Qwen3.8-27B-DFlash2-GGUF (Q8_0). Speedup over plain decode with thinking off, for code / prose / math prompts:

device, target plain tok/s yours ProCreations incoai
RTX 4090, PQ2_0 87.2 1.70 / 1.16 / 2.40 1.75 / 1.27 / 2.37 1.70 / 1.00 / 2.30
RTX 5090, PQ2_0 129.0 1.28 / 0.86 / 1.90 1.30 / 0.91 / 1.73 1.22 / 0.79 / 1.80
RTX 4090, PTQ1_0 92.8 1.58 / 1.16 / 2.19
RTX 5090, PTQ1_0 120.9 1.52 / 1.16 / 2.19
M5 Pro, PQ2_0 27.8 1.03 / 0.52 / 1.19 1.07 / 0.58 / 1.17 1.03 / 0.45 / 1.10
M5 Pro, PTQ1_0 26.5 0.48 / 0.27 / 0.53

Thinking on gives similar numbers on the 4090 and somewhat lower ones on the 5090. Thanks also for the L4 benchmark; the numbers above line up with it: roughly 2x on code and math, much less on prose.

Output. On Metal all 24 speculative runs match plain greedy decode byte for byte. On CUDA 28 of 48 match. Where a run differs, all three drafters diverge at the same character, which fits your explanation that the batched verify rounds differently from single-row decode rather than anything drafter-dependent.

Metal. It gives no speedup there, and PTQ1_0 gets slower. The cause is on our side: the Metal PTQ1_0 mat-vec has no fast path for 2 to 8 columns (17408 x 5120 takes 95 us at one column and 620 us at two, against 176 us for PQ2_0). We'll look at that separately; it isn't a reason to hold this PR.

Upstream follow-ups not included, which could be separate small cherry-picks: cc231cb (NVFP4 scales in DFlash attention), 662a0b0 (fused encoder KV injection), and fa67698 / b0dcb81 (speculation after image input).

@NakliTechie

Copy link
Copy Markdown
Author

Thanks @bri-prism for the careful review and the cross-device numbers.

One more result that needs no code change here: stacking prompt lookup in front of DFlash2 with the existing priority list, --spec-type ngram-mod,draft-dflash (ngram-mod defaults: 24-token match, 48-64-token drafts). When an exact match exists it drafts past the 8-token DFlash2 block; otherwise DFlash2 runs as usual.

NVIDIA L4, llama-server, greedy, one slot, thinking off, target PQ2_0, r3 drafter Q4_K_M. DFlash2 alone and stacked ran on the same machine in the same session; speedup vs plain PTQ1_0:

set DFlash2 stacked change
HumanEval (164) 2.69x 3.58x +33%
code-edit (80, refactor a given function) 2.46x 3.15x +28%
MATH-500 / GSM8K (100 each) 2.20x / 2.17x 2.19x / 2.14x -0.3% / -1.4%
MT-Bench turn 1 / turn 2 (80 each) 1.39x / 1.44x 1.39x / 1.45x 0%
Spec-Bench RAG / summarization (80 each) 1.56x / 1.30x 1.55x / 1.29x -1%
MBPP (100) 2.17x 2.07x -4.5%
  • pass@1 moved by at most 2 problems per set.
  • HumanEval averages 10.2 tokens per verify step with stacking.
  • MBPP loses a little: its prompts include the test asserts, and echoed names occasionally trigger a long match with a wrong continuation.
  • A looser ngram-mod (12-token match, 8-32-token drafts) was worse than the defaults on every set.

On Metal I see the same thing you describe from the small-M side: an in-block copy draft (at most 8 rows) gave no gain in our MLX and WebGPU ports, so the long drafts are what pay, and those need a fast verify above 8 rows.

Rows and summaries: https://huggingface.co/datasets/naklitechie/bonsai2-dflash2-bench (folder stacking/).

@bri-prism

Copy link
Copy Markdown

Windows build check (CI for this PR is stuck in the queue): merged onto prism @ bda59eaa0, merge clean. This is compile + CPU smoke only on Windows 11 / x86-64; the Metal on-device results are in the review from the M5 session.

toolchain config result
MSVC 19.44 (VS 2022 Build Tools) GGML_NATIVE=OFF GGML_BACKEND_DL=ON GGML_CPU_ALL_VARIANTS=ON, tools, server, tests ✅ llama / llama-server / llama-cli / test-backend-ops build, 0 warnings in the files this PR touches
MinGW-w64 GCC 16.2 (MSYS2 UCRT64) GGML_NATIVE=ON ✅ builds. The 6 warnings in touched files are all pre-existing -Wdeprecated-enum-enum-conversion in test-backend-ops.cpp (from ggml-org#18327 / ggml-org#17577 upstream), none from this PR

Smoke test: the MSVC-built llama-server loads a qwen35 2B (PTQ1_0, CPU) and completes "The capital of France is" → " Paris. …".

One unrelated Windows issue came up that is pre-existing on prism, not from this PR: with LLAMA_BUILD_TESTS=ON, MSVC fails to compile tests/test-dspark-forward.cpp. It uses POSIX setenv at L622/633/659 (from 08ac1fd17 "dspark: validate mask selection…"), and MSVC only has _putenv_s. Release builds set LLAMA_BUILD_TESTS=OFF, so they're unaffected, but any Windows CI job that builds tests would fail. It's worth a small #ifdef _WIN32 shim in its own PR.

Not covered: Ubuntu, and running DFlash2 end to end (no DFlash2 drafter here).

Tested with Claude Code.

@bri-prism
bri-prism merged commit 279df66 into PrismML-Eng:prism Sep 25, 2026
3 of 5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants