Repository navigation
glm5-next : keep the indexer compute buffer bounded at long contexts - #243
danielhanchen wants to merge 2 commits into
Conversation
Two things grew the CUDA compute buffer with n_kv * n_ubatch: - build_kpool_select reshaped the host-side kq_mask (and gather_mask, new_pool_idxs) once per DSA layer. The scheduler copies every view of a host input to the device separately, so each layer held its own [n_kv, n_ubatch] f16 copy of the mask. Build those views once per graph. - The reserve re-pools every pool (kv size / kpool * n_seq_max pools), and the pooling kept several f32 [head_dim, kpool] copies of all of them alive. Pool in chunks of 16384 pools; decode-time graphs still use one chunk. GLM-5.3-Flash UD-IQ1_S, -c 244480 --kv-unified -fa on, CUDA0 compute buffer: -np 4 -ub 512: 4923 -> 1153 MiB, -np 1: 3275 -> 941 MiB, -np 4 -ub 2048: 13954 -> 4463 MiB. Perplexity is unchanged.
|
You have reached your Codex usage limits for security reviews. Please try again later. |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: aba4a4c1db
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| // Views of the inputs above, built once per graph: the scheduler copies every view of a host input to | ||
| // the device on its own, so a per-layer reshape of the [n_kv, n_tokens] mask costs a copy per layer. |
There was a problem hiding this comment.
Keep each comment sentence on one line
This comment hard-wraps one sentence across two lines, and the same pattern appears at lines 806-807. Keep each sentence on a single line because the repository explicitly prohibits splitting comment sentences to fit a fixed width.
AGENTS.md reference: AGENTS.md:L74-L81
Useful? React with 👍 / 👎.
# Conflicts: # src/models/glm5-next.cpp
…at merge past b11453, drop unslothai#247, pin unslothai#251 Upstream K2 Horizon (ggml-org#29535), the glm5-next gather removal (ggml-org#30042) and the GLM5-Next MTP graph (ggml-org#29928) broke the Inkling, unslothai#243 and unslothai#241 pins. Each branch now carries a merge of upstream master. unslothai#247 is in every tag from b11452 on. unslothai#251 adds --moe-cache-mib auto on top of ggml-org#29887.
unsloth: repin Inkling, unslothai#241 and unslothai#243 past b11453, drop unslothai#247, pin --moe-cache-mib auto (unslothai#251)
Summary
On GLM-5.3-Flash the glm5-next graph from ggml-org#27773 reserves a much larger CUDA compute buffer than the previous GLM-5-Next pin did: 4939 MiB vs 1635 MiB at
-c 244480 -np 4 --kv-unified -fa on. On a 96 GB card that is the difference between the model plus projector fitting and an out of memory error while loading the mmproj, and it also leaves no room for the NextN MTP head.Two things grew with
n_kv * n_ubatch:build_kpool_selectreshaped the host-sidekq_mask(andgather_mask,new_pool_idxs) once per DSA layer. The scheduler copies every view of a host input to the device separately, so each layer held its own[n_kv, n_ubatch]f16 copy of the mask, about 2.6 GB at 244480 cells. These views are now built once per graph.kv size / kpool * n_seq_maxof them), and the pooling kept several f32[head_dim, kpool]copies of all of them alive at once, about 1.9 GB. Pooling now runs in chunks of 16384 pools. Decode-time graphs re-pool a handful of pools and still use a single chunk.One file,
src/models/glm5-next.cpp, +43/-14. No new flags. Stock upstream master has the same code and the same numbers.Testing
GLM-5.3-Flash UD-IQ1_S, B200,
-c 244480 --kv-unified -fa on, CUDA0 compute buffer:-np 4 -ub 512-np 1 -ub 512-np 4 -ub 2048-np 4-c 4096, 4 chunks) is identical before and after on the mix (4.0045) and on master (4.0103), and also with a test build that forces 7-pool chunks so chunking runs at decode time.--spec-type draft-mtp, and on a 14k-token prompt. MTP acceptance is identical (2360 of 2905).-c 97280 -np 4and the vision projector uses 95488 MiB, so it fits a 97 GB card at about 1.45x the decode speed of no speculation.