Repository navigation
Conversation
|
When you convert the model, try to add llama.cpp/convert_hf_to_gguf.py Lines 10974 to 10981 in 294b2b4 |
|
IIRC they forgot to add |
|
Oh sorry I didn't notice that it's private 😅 temporary closing this to keep it under the radar |
|
I've fixed the Qwen3-VL-Embedding issues in llama.cpp and verified the fix with regression tests. Check out the code here: https://github.com/Tokimorphling/qwen3-vl-embedding The implementation of the Qwen3-VL series in llama.cpp seems to be problematic. |
|
Are there any plans to merge official support for this model? Having a multimodal embedding model would open up a lot of doors with llama.cpp |
|
@ethanmc22 You just need a few extra steps on both the server and client sides: Server side (one env var plus the existing flags): LLAMA_MEDIA_MARKER=<__media__> llama-server \
-m Qwen.Qwen3-VL-Embedding-2B.Q4_K_M.gguf \
--mmproj mmproj-Qwen.Qwen3-VL-Embedding-2B.f16.gguf \
--embedding --pooling last --embd-normalize 2Without Client side, send the same shape {
"input": [
{
"prompt_string": "Describe this image. <__media__>",
"multimodal_data": [
"<base64>"
]
}
]
}I'm using this exact quant: https://huggingface.co/DevQuasar/Qwen.Qwen3-VL-Embedding-2B-GGUF |
|
Thank you so much!! I just tried this out and it works unbelievably well!!! I use local models for mechanical engineering work, a lot of it is done in pdf's, graphs, tables, hand writing, sketches, etc, which cant be converted to text cleanly and needs native vision to work but uploading whole 50 page pdfs as images is far too slow and burns 100k+ tokens. Just tried it out with a quick python script and it works perfectly for retrieval tasks with all of my handwritten notes, lookup graphs/tables, pdfs, etc. Hopefully the qwen3vl reranker gets support soon aswell, i wasnt able to get this working unfortunatly. |
|
Hi @ngxson I've implemented native image + text support for both the Embedding and Reranker endpoints. This was in support of using Qwen3VL-Embedding and Reranker models for RAG workflows. Related to #25921 https://github.com/timothywang21/llama.cpp-Qwen3VL-Support Changes I made
Let me know how you want to proceed. |
…ing) ## Overview Right now, multimodal support for embeddings is possible but requires janky workarounds like ggml-org#18665 (comment) This PR consolidates these workarounds by implementing native multimodal support as well as introducing an OpenAI-style API for the /v1/embeddings endpoint. This new OAI-style API is the defacto standard that is used by Openrouter and other providers for multimodal embeddings. The previous legacy API is still supported. **Legacy API (still supported):** ```json { "input": { "prompt_string": "Describe this image<__media__>", "multimodal_data": ["iVBORw0KGgoAAAANSUhEUg..."] } } ``` **OAI-compatible API (new), multiple content arrays in one request:** ```json { "input": [ { "content": [ { "type": "image_url", "image_url": { "url": "data:image/jpeg;base64,/9j/4AAQSkZJRg..." } }, { "type": "text", "text": "Describe this image" } ] }, { "content": [ { "type": "text", "text": "This is a second content array you can pass in one request" } ] } ] } ``` **Incorrect (bare content array - rejected with 400):** ```json { "input": [ { "type": "image_url", "image_url": { "url": "data:image/jpeg;base64,/9j/4AAQSkZJRg..." } }, { "type": "text", "text": "Describe this image" } ] } ``` Assisted-by: Opencode Qwen3.8 27B
Target support: https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B
Important
the original Qwen3-VL-Embedding model is missing 1_Pooling, I don't think it's actually ready to be used unless Qwen team fixed it (I already reached out to them, but got no responses)
But currently, the model is missing
1_Pooling, so it cannot be correctly converted to GGUFThis PR aims to support mixed text+image (and maybe audio input for models supporting it) using OAI-compat
content-like schema:{ "input": [ { "type": "text", "text": "mixed text and image input" }, { "type": "image", "image_url": { "url": "https://huggingface.co/ggml-org/tinygemma3-GGUF/resolve/main/test/11_truck.png" } } ] }