Skip to content

glm5-next : keep the indexer compute buffer bounded at long contexts - #243

Open
danielhanchen wants to merge 2 commits into
base/upstream-5ad1c5da0from
glm5-next-compute-buffer
Open

danielhanchen wants to merge 2 commits into
base/upstream-5ad1c5da0from
glm5-next-compute-buffer

Conversation

@danielhanchen

Copy link
Copy Markdown
Member

Summary

On GLM-5.3-Flash the glm5-next graph from ggml-org#27773 reserves a much larger CUDA compute buffer than the previous GLM-5-Next pin did: 4939 MiB vs 1635 MiB at -c 244480 -np 4 --kv-unified -fa on. On a 96 GB card that is the difference between the model plus projector fitting and an out of memory error while loading the mmproj, and it also leaves no room for the NextN MTP head.

Two things grew with n_kv * n_ubatch:

  • build_kpool_select reshaped the host-side kq_mask (and gather_mask, new_pool_idxs) once per DSA layer. The scheduler copies every view of a host input to the device separately, so each layer held its own [n_kv, n_ubatch] f16 copy of the mask, about 2.6 GB at 244480 cells. These views are now built once per graph.
  • The full-context reserve re-pools every pool (kv size / kpool * n_seq_max of them), and the pooling kept several f32 [head_dim, kpool] copies of all of them alive at once, about 1.9 GB. Pooling now runs in chunks of 16384 pools. Decode-time graphs re-pool a handful of pools and still use a single chunk.

One file, src/models/glm5-next.cpp, +43/-14. No new flags. Stock upstream master has the same code and the same numbers.

Testing

GLM-5.3-Flash UD-IQ1_S, B200, -c 244480 --kv-unified -fa on, CUDA0 compute buffer:

before after
Unsloth mix, -np 4 -ub 512 4923 MiB 1153 MiB
Unsloth mix, -np 1 -ub 512 3275 MiB 941 MiB
Unsloth mix, -np 4 -ub 2048 13954 MiB 4463 MiB
upstream master 50569eb, -np 4 4939 MiB 1161 MiB
  • Perplexity (wikitext, -c 4096, 4 chunks) is identical before and after on the mix (4.0045) and on master (4.0103), and also with a test build that forces 7-pool chunks so chunking runs at decode time.
  • Greedy output is byte-identical before and after, with and without --spec-type draft-mtp, and on a 14k-token prompt. MTP acceptance is identical (2360 of 2905).
  • Decode speed, interleaved runs on one GPU: 68.2 / 68.0 tok/s before vs 70.5 / 69.2 after without speculation; 104.4 / 103.5 vs 105.9 / 109.3 with MTP.
  • With this change, Studio's GLM-5.3-Flash launch with the MTP head at -c 97280 -np 4 and the vision projector uses 95488 MiB, so it fits a 97 GB card at about 1.45x the decode speed of no speculation.
  • The commit cherry-picks cleanly onto this base and three-way merges onto current upstream master.

Two things grew the CUDA compute buffer with n_kv * n_ubatch:

- build_kpool_select reshaped the host-side kq_mask (and gather_mask,
  new_pool_idxs) once per DSA layer. The scheduler copies every view of a
  host input to the device separately, so each layer held its own
  [n_kv, n_ubatch] f16 copy of the mask. Build those views once per graph.
- The reserve re-pools every pool (kv size / kpool * n_seq_max pools), and the
  pooling kept several f32 [head_dim, kpool] copies of all of them alive.
  Pool in chunks of 16384 pools; decode-time graphs still use one chunk.

GLM-5.3-Flash UD-IQ1_S, -c 244480 --kv-unified -fa on, CUDA0 compute buffer:
-np 4 -ub 512: 4923 -> 1153 MiB, -np 1: 3275 -> 941 MiB,
-np 4 -ub 2048: 13954 -> 4463 MiB. Perplexity is unchanged.
@danielhanchen
danielhanchen requested a review from CISC as a code owner October 6, 2026 06:44
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for security reviews. Please try again later.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-06T06:47:59.460527Z aba4a4c PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: aba4a4c1db

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/models/glm5-next.cpp
Comment on lines +281 to +282
// Views of the inputs above, built once per graph: the scheduler copies every view of a host input to
// the device on its own, so a per-layer reshape of the [n_kv, n_tokens] mask costs a copy per layer.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep each comment sentence on one line

This comment hard-wraps one sentence across two lines, and the same pattern appears at lines 806-807. Keep each sentence on a single line because the repository explicitly prohibits splitting comment sentences to fit a fixed width.

AGENTS.md reference: AGENTS.md:L74-L81

Useful? React with 👍 / 👎.

# Conflicts:
#	src/models/glm5-next.cpp
@danielhanchen
danielhanchen requested a review from ngxson as a code owner October 7, 2026 14:00
EmeraldBitTwizzler pushed a commit to EmeraldBitTwizzler/llama.cpp that referenced this pull request Oct 9, 2026
…at merge past b11453, drop unslothai#247, pin unslothai#251

Upstream K2 Horizon (ggml-org#29535), the glm5-next gather removal (ggml-org#30042) and the
GLM5-Next MTP graph (ggml-org#29928) broke the Inkling, unslothai#243 and unslothai#241 pins. Each
branch now carries a merge of upstream master. unslothai#247 is in every tag from
b11452 on. unslothai#251 adds --moe-cache-mib auto on top of ggml-org#29887.
EmeraldBitTwizzler pushed a commit to EmeraldBitTwizzler/llama.cpp that referenced this pull request Oct 9, 2026
@danielhanchen
danielhanchen changed the base branch from base/upstream-1fb7ef3e3 to base/upstream-5ad1c5da0 October 9, 2026 12:10

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant