Skip to content

cuda: top-k MoE should always fire - #28432

Merged
ggerganov merged 1 commit into
masterfrom
aman/top-k-alloc-dep
Sep 23, 2026
Merged

ggerganov merged 1 commit into
masterfrom
aman/top-k-alloc-dep

Conversation

@am17an

@am17an am17an commented Sep 5, 2026

Copy link
Copy Markdown
Contributor

Overview

Cont #27301. Allow top-k moe to always fire.

Additional information

Requirements

@am17an
am17an requested a review from a team as a code owner September 5, 2026 09:09
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Sep 5, 2026
@am17an

am17an commented Sep 11, 2026

Copy link
Copy Markdown
Contributor Author

@ggml-org/ggml-cuda can someone take a look here?

@IMbackK IMbackK self-assigned this Sep 11, 2026
@JohannesGaessler JohannesGaessler self-assigned this Sep 20, 2026

@JohannesGaessler JohannesGaessler left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not that knowledgeable when it comes to the current fusion code but these changes seem correct to me.

@am17an am17an added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Sep 23, 2026
@ggerganov
ggerganov merged commit 1a67982 into master Sep 23, 2026
22 of 24 checks passed
@BrewTestBot BrewTestBot mentioned this pull request Sep 23, 2026
1 task done
@ggerganov

Copy link
Copy Markdown
Member

@am17an

am17an commented Sep 23, 2026

Copy link
Copy Markdown
Contributor Author

@ggerganov this is a GCC bug. I think a reserve might get rid of the warning. Will put up a PR

iki007 added a commit to iki007/llama.cpp that referenced this pull request Sep 23, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
smalinin pushed a commit to smalinin/llama.cpp that referenced this pull request Sep 23, 2026
@ggerganov
ggerganov deleted the aman/top-k-alloc-dep branch September 24, 2026 06:15
iki007 added a commit to iki007/llama.cpp that referenced this pull request Sep 24, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Sep 24, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Sep 25, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
mclrcha added a commit to mclrcha/llama.cpp that referenced this pull request Sep 25, 2026
… alloc deps

Upstream ggml-org#28432 adds allocator dependencies so the fused top-k MoE router
always fires. On Qwen3.6-35B-A3B this changes outputs (f32 reduction order,
near-tie expert choices flip; KLD 0.023 vs unfused, same precision) and gives
+1% prefill. Default stays on; =0 reproduces the pre-sync outputs bit for bit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
mclrcha added a commit to mclrcha/llama.cpp that referenced this pull request Sep 25, 2026
… alloc deps

Upstream ggml-org#28432 adds allocator dependencies so the fused top-k MoE router
always fires. On Qwen3.6-35B-A3B this changes outputs (f32 reduction order,
near-tie expert choices flip; KLD 0.023 vs unfused, same precision) and gives
+1% prefill. Default stays on; =0 reproduces the pre-sync outputs bit for bit.
iki007 added a commit to iki007/llama.cpp that referenced this pull request Sep 26, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Sep 26, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Sep 27, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Sep 27, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Sep 28, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
baykalokandemir added a commit to baykalokandemir/llama.cpp-bells-tortured that referenced this pull request Sep 28, 2026
…ggml-org#28118 and ggml-org#28751 do not apply (MTP uses bounded rollback, no causal toggles)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018co7fcJoEQXMxLezg7HpMU
iki007 added a commit to iki007/llama.cpp that referenced this pull request Sep 29, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Sep 30, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Oct 1, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Oct 1, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 1, 2026
iki007 added a commit to iki007/llama.cpp that referenced this pull request Oct 2, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Oct 3, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 4, 2026
LadislavSopko pushed a commit to 0ics-srls/llama.cpp that referenced this pull request Oct 5, 2026
iki007 added a commit to iki007/llama.cpp that referenced this pull request Oct 5, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
iki007 added a commit to iki007/llama.cpp that referenced this pull request Oct 7, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Oct 9, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Oct 10, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Oct 10, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Oct 10, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
iki007 added a commit to iki007/llama.cpp that referenced this pull request Oct 10, 2026
…ys fires

Port of CUDA ggml-org#28432. On multi-row batches the graph allocator could place the
router outputs over the freed logits, and the memory-range check then declined
the fused kernel (78 of 200 routers on Qwen3.6-35B-A3B). graph_optimize now
keeps every router node and its sources allocated until the router's last node.

Arc Pro B70, Qwen3.6-35B-A3B Q4_K_M: pp4 +2.45%, pp2048 +0.74%, tg unchanged,
compute buffer unchanged.

Assisted-by: Claude Opus 5.5
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants