Repository navigation
convert: add MiMo-V2.6 support - #29257
Conversation
Hoist the K3 mxfp4 conversion repack into base.py so it can be reused Remove decoder from mmproj convert
There was a problem hiding this comment.
AFAICT there is an issue with the chat template - tool calls don't work correctly. There is a suggested fix here that on first look does the job: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B/discussions/6. But more investigation is needed.
Also, not yet sure how to convert the speculative model. Is it MTP or DFlash? Or both? Not clear from the source repo.
In any case, this change seems to work to get the base conversion going: https://huggingface.co/ggml-org/MiMo-V2.6-Flash-RL-GGUF
|
We should add a negative branch to the check to make MiMo take the autoparser branch instead of the (incorrect) Qwen branch, I'll work on it. |
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
|
@ggerganov merge ready? |
|
Nope, needs my PR merged: AesSedai#1, unless we merge this and then I quickly pit it up (might be faster since @AesSedai went to sleep I think). |
Just commit it directly. |
|
Can't, no write access to the PR, I tried. |
You should be able to commit directly to |
|
All right, no idea what happened but when using the GitHub Workspace, it wouldn't let me push there. Natively from shell it worked fine. |
|
@CISC can I haz reapprove? |
* convert: add MiMo-V2.6 support Hoist the K3 mxfp4 conversion repack into base.py so it can be reused Remove decoder from mmproj convert * Update conversion/mimo.py * fix: use autoparser --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> Co-authored-by: Piotr Wilkin <piotr.wilkin@syndatis.com>
* convert: add MiMo-V2.6 support Hoist the K3 mxfp4 conversion repack into base.py so it can be reused Remove decoder from mmproj convert * Update conversion/mimo.py * fix: use autoparser --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> Co-authored-by: Piotr Wilkin <piotr.wilkin@syndatis.com>
* convert: add MiMo-V2.6 support Hoist the K3 mxfp4 conversion repack into base.py so it can be reused Remove decoder from mmproj convert * Update conversion/mimo.py * fix: use autoparser --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> Co-authored-by: Piotr Wilkin <piotr.wilkin@syndatis.com>
* convert: add MiMo-V2.6 support Hoist the K3 mxfp4 conversion repack into base.py so it can be reused Remove decoder from mmproj convert * Update conversion/mimo.py * fix: use autoparser --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> Co-authored-by: Piotr Wilkin <piotr.wilkin@syndatis.com> (cherry picked from commit bfd73a8)
* convert: add MiMo-V2.6 support Hoist the K3 mxfp4 conversion repack into base.py so it can be reused Remove decoder from mmproj convert * Update conversion/mimo.py * fix: use autoparser --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co> Co-authored-by: Piotr Wilkin <piotr.wilkin@syndatis.com> (cherry picked from commit bfd73a8)
Overview
This PR adds conversion support for MiMo-V2.6 Pro and Flash. Both of these models use mxfp4 experts like DSv4 and Kimi-K3 do, so I hoisted the K3 repack to
base.pyso it can be re-used. There's also an audio decoder with the mmproj, so I excluded that from the conversion.Additional information
Both models convert correctly and load into llama.cpp with vision. Runtime changes weren't necessary to support the architectures once converted.
Requirements