Repository navigation
Qwen3VLChatHandler gives wrong image dimensions/spatial understanding vs upstream llama-server (same model/mmproj) #152
Description
Activity
No, you've got it wrong;
media_markerinteracts with the underlyingllama.cppto recognize the renderedimage_url—it is not sensitive to<image_pad>.The final point is also incorrect: if the cache doesn't match, re-computation is naturally required; you cannot simply reuse the previous result. If reuse is desired, the program itself must manage the caching of the conversational context.
Furthermore, asking a large model about image resolution is inherently prone to hallucinations; the underlying system performs preprocessing—such as resizing and cropping—on the images anyway. Therefore, using this method to evaluate image size is inappropriate and meaningless.
Thank you for the clarification and for taking the time to look into this — appreciated.
I understand the internal preprocessing (resize/crop) explains why asking the model directly for image resolution is unreliable, and I agree that's expected model behavior rather than a binding bug per se.
That said, I want to push back gently on the practical implication: for a very common real-world use case — get a bounding box from the model, then apply it as a mask/crop back onto the original image (e.g. for inpainting, redaction, OCR-driven editing) — if the model's coordinate output is relative to an internal resized/cropped representation that the caller has no visibility into (no returned scale factor, offset, or reference resolution), then there's no reliable way to map that box back onto the original image. We tested this directly (synthetic images with known ground-truth text position, at multiple sizes/aspect ratios) and the reprojected boxes are off by anywhere from ~10px to 150px depending on how far the image deviates from a square, "typically sized" input — with no fixed offset or scale factor that would let a caller correct for it after the fact.
So while I take the point that asking for raw pixel dimensions is the wrong ask, the same underlying issue makes grounding-for-masking unreliable too, and that's a pretty central use case for a vision chat handler. If there's no way for the binding to expose the actual preprocessing parameters (target size / padding / crop offset) used for a given image, that seems worth documenting clearly so people don't hit this the hard way like we did. Happy to share the full test data if it's useful for anyone looking into this further.
Thanks again for llama-cpp-python — really useful project overall.
The performance of grounding relies heavily on the data used to train the model; the underlying CLIP model typically specifies the resolutions used during training. For instance, the GraniteDocling model achieves greater accuracy in pinpointing the locations of annotations on input images sized at 512x512.
Following up on this — there's now hard evidence for the residual spatial error @sucubus666 raised, which was dismissed above as expected model hallucination.
A separate PR against upstream llama.cpp (ggml-org/llama.cpp#25781, closed but not merged) identifies the actual cause:
resize_position_embeddings()intools/mtmd/models/qwen3vl.cppinterpolates the learned 48×48 vision position embedding grid usingGGML_SCALE_MODE_BILINEAR | GGML_SCALE_FLAG_ANTIALIAS(align_corners=False), while the referencetransformersimplementation (fast_pos_embed_interpolate) uses align_corners=True (torch.linspace(0, side-1, T)). The two conventions differ by a per-axis scale about the image center ofside·(T-1)/(T·(side-1)), growing with patch-grid size and differing per axis for non-square images — which matches exactly the size-dependent, non-square-asymmetric residual error reported here, on top of the 0-1000 rescale from #16880.The PR author verified against a fine-tuned Qwen3-VL emitting per-word boxes at a fixed 1792×2560 input: llama.cpp before the fix matches
transformersforced toalign_corners=Falsebin-for-bin, and after the fix matches the realalign_corners=Truereference within bf16 noise — across ~220 boxes, fitting the pre-fix scale factor to 1.0126(x)/1.0134(y), matching the analytic prediction 1.0122/1.0149.Given this, would you consider adopting the align_corners fix from that PR into this fork's
Qwen3VLChatHandler/ vendored llama.cpp? Happy to test a build if useful.
Qwen3VL-8B-Instruct-Q4_K_M.gguf) + same mmproj (mmproj-Qwen3VL-8B-Instruct-F16.gguf), same image (1920x1080 PNG).llama-server(native,--jinja, using the model's embedded chat template) → asked "what are the exact pixel dimensions of this image?" → correct answer: 1920x1080.llama-cpp-python(this fork, v0.3.42) usingQwen3VLChatHandlerdirectly in Python, same image, same question → wrong answer, inconsistent across runs (768x451, then 1280x720 on a re-run).chunk_n_tokens: 2040≈ 1920x1080 area / patch_area).Qwen3VLChatHandler's built-in Jinja chat template (hardcoded inllama_multimodal.py) uses<__media__>as the image placeholder, while the model's own embeddedtokenizer.chat_template(used byllama-server --jinja) uses<|image_pad|>. Wondering if this divergence affects how image tokens are positioned/interpreted (rope_2d / spatial reasoning) even though the raw token count matches.Llama.longest_token_prefix [Fast Exit 2]: First token mismatch... Prefix mismatch. Truncating KV cache from 189 to 0even with a fresh/reset session — possibly related to Misc. bug: Qwen3-VL on llama-server fails on second request (KV Cache Bug) & /slots/reset returns 501 ggml-org/llama.cpp#17200.Happy to provide the full repro script and logs.