Repository navigation
cuda: top-k MoE should always fire - #28432
Merged
Merged
Conversation
Contributor
Author
|
@ggml-org/ggml-cuda can someone take a look here? |
JohannesGaessler
approved these changes
Sep 22, 2026
JohannesGaessler
left a comment
Contributor
There was a problem hiding this comment.
I'm not that knowledgeable when it comes to the current fusion code but these changes seem correct to me.
1 task done
Member
|
@am17an This change breaks builds with nvcc 12.4: |
Contributor
Author
|
@ggerganov this is a GCC bug. I think a |
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Sep 23, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
smalinin
pushed a commit
to smalinin/llama.cpp
that referenced
this pull request
Sep 23, 2026
(cherry picked from commit 1a67982)
4 tasks
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Sep 24, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Sep 24, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Sep 25, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
mclrcha
added a commit
to mclrcha/llama.cpp
that referenced
this pull request
Sep 25, 2026
… alloc deps Upstream ggml-org#28432 adds allocator dependencies so the fused top-k MoE router always fires. On Qwen3.6-35B-A3B this changes outputs (f32 reduction order, near-tie expert choices flip; KLD 0.023 vs unfused, same precision) and gives +1% prefill. Default stays on; =0 reproduces the pre-sync outputs bit for bit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
mclrcha
added a commit
to mclrcha/llama.cpp
that referenced
this pull request
Sep 25, 2026
… alloc deps Upstream ggml-org#28432 adds allocator dependencies so the fused top-k MoE router always fires. On Qwen3.6-35B-A3B this changes outputs (f32 reduction order, near-tie expert choices flip; KLD 0.023 vs unfused, same precision) and gives +1% prefill. Default stays on; =0 reproduces the pre-sync outputs bit for bit.
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Sep 26, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Sep 26, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Sep 27, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Sep 27, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Sep 28, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
baykalokandemir
added a commit
to baykalokandemir/llama.cpp-bells-tortured
that referenced
this pull request
Sep 28, 2026
…ggml-org#28118 and ggml-org#28751 do not apply (MTP uses bounded rollback, no causal toggles) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018co7fcJoEQXMxLezg7HpMU
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Sep 29, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Sep 30, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Oct 1, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Oct 1, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
edwardyoon
pushed a commit
to edwardyoon/focus-llama
that referenced
this pull request
Oct 1, 2026
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Oct 2, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Oct 3, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
edwardyoon
pushed a commit
to edwardyoon/focus-llama
that referenced
this pull request
Oct 4, 2026
(cherry picked from commit 1a67982)
LadislavSopko
pushed a commit
to 0ics-srls/llama.cpp
that referenced
this pull request
Oct 5, 2026
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Oct 5, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
frostyautumnleaf
pushed a commit
to frostyautumnleaf/llama.cpp
that referenced
this pull request
Oct 5, 2026
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Oct 7, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Oct 9, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Oct 10, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Oct 10, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Oct 10, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
iki007
added a commit
to iki007/llama.cpp
that referenced
this pull request
Oct 10, 2026
…ys fires Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the router outputs over the freed logits, and the memory-range check then declined the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now keeps every router node and its sources allocated until the router's last node. Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged, compute buffer unchanged. Assisted-by: Claude Opus 5.5
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Cont #27301. Allow top-k moe to always fire.
Additional information
Requirements