Skip to content

Fix memory allocation for MTP layers for avoid cudaMalloc error - #25574

Closed
smalinin wants to merge 0 commit into
ggml-org:masterfrom
smalinin:master
Closed

smalinin wants to merge 0 commit into
ggml-org:masterfrom
smalinin:master

Conversation

@smalinin

Copy link
Copy Markdown
Contributor
  • include NextN/MTP layers in auto VRAM fitting
  • preserve TENSOR_SKIP for fused QKV tensors

Overview

Fix issue with memory allocation, when MTP model is loaded and --n-cpu-moe=0 options is used.
I tried to load MiMo-V2.5 model and got the next error:

[39535] 0.00.065.618 I cmn  common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
[39535] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model"],"current":"text_model","value":0.0}}
[39535] 0.00.522.196 I srv    load_model: loading model '/home/.models/AesSedai/MiMo-V2.5-GGUF/MiMo-V2.5-Q4_K_M-00001-of-00005.gguf'
[39535] 0.04.650.599 W load: control-looking token: 128247 '</s>' was not control-type; this is probably a bug in the model. its type will be overridden
[39535] 0.04.689.840 W model has unused tensor blk.48.attn_output.weight (size = 35651584 bytes) -- ignoring
[39535] 0.04.689.843 W model has unused tensor blk.48.attn_norm.weight (size = 16384 bytes) -- ignoring
[39535] 0.04.689.844 W model has unused tensor blk.48.attn_sinks.weight (size = 256 bytes) -- ignoring
[39535] 0.04.689.845 W model has unused tensor blk.48.ffn_norm.weight (size = 16384 bytes) -- ignoring
...
...
[39535] 0.04.689.928 W model has unused tensor blk.50.nextn.eh_proj.weight (size = 35651584 bytes) -- ignoring
[39535] 0.04.689.929 W model has unused tensor blk.50.nextn.enorm.weight (size = 16384 bytes) -- ignoring
[39535] 0.04.689.930 W model has unused tensor blk.50.nextn.hnorm.weight (size = 16384 bytes) -- ignoring
[39535] 0.04.689.931 W model has unused tensor blk.50.layer_output_norm.weight (size = 16384 bytes) -- ignoring
[39535] 0.04.691.073 E ggml_backend_cuda_buffer_type_alloc_buffer: allocating 21456.63 MiB on device 0: cudaMalloc failed: out of memory
[39535] 0.04.691.076 E alloc_tensor_range: failed to allocate CUDA0 buffer of size 22498909440
[39535] 0.04.709.937 E llama_model_load: error loading model: unable to allocate CUDA0 buffer
[39535] 0.04.709.947 E llama_model_load_from_file_impl: failed to load model
[39535] 0.04.709.951 E cmn  common_init_: failed to load model '/home/.models/AesSedai/MiMo-V2.5-GGUF/MiMo-V2.5-Q4_K_M-00001-of-00005.gguf'
[39535] 0.04.709.955 E srv    load_model: failed to load model, '/home/.models/AesSedai/MiMo-V2.5-GGUF/MiMo-V2.5-Q4_K_M-00001-of-00005.gguf'
[39535] 0.04.709.956 I srv    operator(): operator(): cleaning up before exit...
[39535] 0.04.712.567 E srv  llama_server: exiting due to model loading error
0.11.854.290 I srv    operator(): instance name=MiMo-V2.5-Q4_K_M=128=image=AES exited with status 1

After the fix, everything is loaded without errors:

[47093] 0.00.068.110 I cmn  common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
[47093] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model"],"current":"text_model","value":0.0}}
[47093] 0.00.498.106 I srv    load_model: loading model '/home/.models/AesSedai/MiMo-V2.5-GGUF/MiMo-V2.5-Q4_K_M-00001-of-00005.gguf'
[47093] 0.04.378.963 W load: control-looking token: 128247 '</s>' was not control-type; this is probably a bug in the model. its type will be overridden
[47093] 0.04.426.445 W model has unused tensor blk.48.attn_qkv.weight (size = 64618496 bytes) -- ignoring
[47093] 0.04.426.454 W model has unused tensor blk.48.attn_output.weight (size = 35651584 bytes) -- ignoring
[47093] 0.04.426.455 W model has unused tensor blk.48.attn_norm.weight (size = 16384 bytes) -- ignoring
[47093] 0.04.426.456 W model has unused tensor blk.48.attn_sinks.weight (size = 256 bytes) -- ignoring
[47093] 0.04.426.457 W model has unused tensor blk.48.ffn_norm.weight (size = 16384 bytes) -- ignoring
[47093] 0.04.426.458 W model has unused tensor blk.48.ffn_gate.weight (size = 71303168 bytes) -- ignoring
[47093] 0.04.426.459 W model has unused tensor blk.48.ffn_down.weight (size = 71303168 bytes) -- ignoring
...
[47093] 0.04.426.504 W model has unused tensor blk.50.ffn_down.weight (size = 71303168 bytes) -- ignoring
[47093] 0.04.426.505 W model has unused tensor blk.50.ffn_up.weight (size = 71303168 bytes) -- ignoring
[47093] 0.04.426.511 W model has unused tensor blk.50.nextn.eh_proj.weight (size = 35651584 bytes) -- ignoring
[47093] 0.04.426.512 W model has unused tensor blk.50.nextn.enorm.weight (size = 16384 bytes) -- ignoring
[47093] 0.04.426.513 W model has unused tensor blk.50.nextn.hnorm.weight (size = 16384 bytes) -- ignoring
[47093] 0.04.426.514 W model has unused tensor blk.50.layer_output_norm.weight (size = 16384 bytes) -- ignoring
[47093] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model"],"current":"text_model","value":0.0}}
[47093] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model"],"current":"text_model","value":0.0012509445659816265}}
[47093] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model"],"current":"text_model","value":0.008544351905584335}}
[47093] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model"],"current":"text_model","value":0.01632601208984852}}
...
[47093] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model"],"current":"text_model","value":0.9843482375144958}}
[47093] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model"],"current":"text_model","value":0.9936332106590271}}
[47093] cmd_child_to_router:state:{"state":"loading","payload":{"stages":["text_model"],"current":"text_model","value":1.0}}
[47093] 3.42.892.845 I srv    load_model: initializing, n_slots = 1, n_ctx_slot = 128000, kv_unified = 'false'
[47093] 3.42.910.322 I srv          init: chat template supports preserving reasoning, consider enabling it via --reasoning-preserve
[47093] 3.42.910.356 I srv  llama_server: model loaded
[47093] 3.42.910.359 I srv  llama_server: listening on http://127.0.0.1:47093
[47093] cmd_child_to_router:state:{"state":"ready","payload":{"id":"MiMo-V2.5-Q4_K_M=128=image=AES","aliases":["MiMo-V2.5-Q4_K_M=128=image=AES"],"tags":[],"object":"model","created":1783777142,"owned_by":"llamacpp","meta":{"vocab_type":2,"n_vocab":152576,"n_ctx":128000,"n_ctx_train":1048576,"n_embd":4096,"n_params":309766601088,"size":190777255424,"ftype":"Q8_0"}}}
3.48.582.163 I srv  proxy_reques: proxying request to model MiMo-V2.5-Q4_K_M=128=image=AES on port 47093

The problem was because:

  • --fit didn't use Next/MTP layers
  • TENSOR_SKIP was lost in create_tensor_qkv()

Additional information

Tested with model https://huggingface.co/AesSedai/MiMo-V2.5-GGUF Q4_K_M
with next options:

[*]
flash-attn = 1 
batch-size = 4096
ubatch-size = 4096
parallel = 1
threads = 20

[MiMo-V2.5-Q4_K_M=128=image=AES]
model = /home/.models/AesSedai/MiMo-V2.5-GGUF/MiMo-V2.5-Q4_K_M-00001-of-00005.gguf
no-mmproj = true
ctx-size = 128000
n-cpu-moe = 0
repeat-penalty = 1.03
temp = 0.8
top-p = 0.95
top-k = 0
min-p = 0.0
cache-type-k = q8_0 
cache-type-v = q8_0
reasoning-budget = 4096
no-mmap = true

Hardware:
4 GPU's with VRAM=140Gb of memory in total
CPU RAM = 220Gb

Requirements

  1. Learning, exploration, and understanding the codebase
  2. rechecking fix code

@mattzink

Copy link
Copy Markdown

I can validate that this fixed my error switching between non-MOE MTP models. Thanks!!

Comment thread common/fit.cpp
@CISC

CISC commented Aug 4, 2026

Copy link
Copy Markdown
Member

Rebase and we can merge after CI.

@github-actions github-actions Bot added documentation Improvements or additions to documentation model Model specific build Compilation issues testing Everything test related Vulkan Issues specific to the Vulkan backend devops improvements to build systems and github actions labels Aug 4, 2026
@github-actions github-actions Bot added server ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language Apple Metal https://en.wikipedia.org/wiki/Metal_(API) OpenCL Issues specific to the OpenCL backend Hexagon mtmd Related to multimodal functionality (video/image/audio) CUDA Related to the CUDA backend jinja parser Issues related to the jinja parser AMD ZenDNN Issues related to the AMD ZenDNN backend OpenVINO WebGPU server/ui conversion vendor labels Aug 4, 2026
@smalinin smalinin closed this Aug 4, 2026
@smalinin

smalinin commented Aug 4, 2026

Copy link
Copy Markdown
Contributor Author

@CISC
Sorry, something was wrong with rebase, so I was need to recreate pull request
New created #26605

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

AMD ZenDNN Issues related to the AMD ZenDNN backend Apple Metal https://en.wikipedia.org/wiki/Metal_(API) build Compilation issues conversion CUDA Related to the CUDA backend devops improvements to build systems and github actions documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning Hexagon jinja parser Issues related to the jinja parser model Model specific mtmd Related to multimodal functionality (video/image/audio) OpenCL Issues specific to the OpenCL backend OpenVINO server/ui server SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language testing Everything test related vendor Vulkan Issues specific to the Vulkan backend WebGPU

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants