Skip to content

Qwen3VLChatHandler gives wrong image dimensions/spatial understanding vs upstream llama-server (same model/mmproj) #152

Description

@sucubus666
  • Same model (Qwen3VL-8B-Instruct-Q4_K_M.gguf) + same mmproj (mmproj-Qwen3VL-8B-Instruct-F16.gguf), same image (1920x1080 PNG).
  • Via llama-server (native, --jinja, using the model's embedded chat template) → asked "what are the exact pixel dimensions of this image?" → correct answer: 1920x1080.
  • Via llama-cpp-python (this fork, v0.3.42) using Qwen3VLChatHandler directly in Python, same image, same question → wrong answer, inconsistent across runs (768x451, then 1280x720 on a re-run).
  • Confirmed the image tensor itself is correct size (1920,1080) right before encoding, and confirmed image token count is correct (chunk_n_tokens: 2040 ≈ 1920x1080 area / patch_area).
  • Noticed Qwen3VLChatHandler's built-in Jinja chat template (hardcoded in llama_multimodal.py) uses <__media__> as the image placeholder, while the model's own embedded tokenizer.chat_template (used by llama-server --jinja) uses <|image_pad|>. Wondering if this divergence affects how image tokens are positioned/interpreted (rope_2d / spatial reasoning) even though the raw token count matches.
  • Also observed in logs: Llama.longest_token_prefix [Fast Exit 2]: First token mismatch... Prefix mismatch. Truncating KV cache from 189 to 0 even with a fresh/reset session — possibly related to Misc. bug: Qwen3-VL on llama-server fails on second request (KV Cache Bug) & /slots/reset returns 501 ggml-org/llama.cpp#17200.

Happy to provide the full repro script and logs.

Activity

  1. JamePeng commented on Jul 14, 2026

    @JamePeng
    Owner

    No, you've got it wrong; media_marker interacts with the underlying llama.cpp to recognize the rendered image_url—it is not sensitive to <image_pad>.

    The final point is also incorrect: if the cache doesn't match, re-computation is naturally required; you cannot simply reuse the previous result. If reuse is desired, the program itself must manage the caching of the conversational context.

    Furthermore, asking a large model about image resolution is inherently prone to hallucinations; the underlying system performs preprocessing—such as resizing and cropping—on the images anyway. Therefore, using this method to evaluate image size is inappropriate and meaningless.

  2. sucubus666 commented on Jul 14, 2026

    @sucubus666
    Author

    Thank you for the clarification and for taking the time to look into this — appreciated.

    I understand the internal preprocessing (resize/crop) explains why asking the model directly for image resolution is unreliable, and I agree that's expected model behavior rather than a binding bug per se.

    That said, I want to push back gently on the practical implication: for a very common real-world use case — get a bounding box from the model, then apply it as a mask/crop back onto the original image (e.g. for inpainting, redaction, OCR-driven editing) — if the model's coordinate output is relative to an internal resized/cropped representation that the caller has no visibility into (no returned scale factor, offset, or reference resolution), then there's no reliable way to map that box back onto the original image. We tested this directly (synthetic images with known ground-truth text position, at multiple sizes/aspect ratios) and the reprojected boxes are off by anywhere from ~10px to 150px depending on how far the image deviates from a square, "typically sized" input — with no fixed offset or scale factor that would let a caller correct for it after the fact.

    So while I take the point that asking for raw pixel dimensions is the wrong ask, the same underlying issue makes grounding-for-masking unreliable too, and that's a pretty central use case for a vision chat handler. If there's no way for the binding to expose the actual preprocessing parameters (target size / padding / crop offset) used for a given image, that seems worth documenting clearly so people don't hit this the hard way like we did. Happy to share the full test data if it's useful for anyone looking into this further.

    Thanks again for llama-cpp-python — really useful project overall.

  3. JamePeng commented on Jul 14, 2026

    @JamePeng
    Owner

    The performance of grounding relies heavily on the data used to train the model; the underlying CLIP model typically specifies the resolutions used during training. For instance, the GraniteDocling model achieves greater accuracy in pinpointing the locations of annotations on input images sized at 512x512.

  4. sucubus666 commented on Jul 16, 2026

    @sucubus666
    Author

    Following up on this — there's now hard evidence for the residual spatial error @sucubus666 raised, which was dismissed above as expected model hallucination.

    A separate PR against upstream llama.cpp (ggml-org/llama.cpp#25781, closed but not merged) identifies the actual cause: resize_position_embeddings() in tools/mtmd/models/qwen3vl.cpp interpolates the learned 48×48 vision position embedding grid using GGML_SCALE_MODE_BILINEAR | GGML_SCALE_FLAG_ANTIALIAS (align_corners=False), while the reference transformers implementation (fast_pos_embed_interpolate) uses align_corners=True (torch.linspace(0, side-1, T)). The two conventions differ by a per-axis scale about the image center of side·(T-1)/(T·(side-1)), growing with patch-grid size and differing per axis for non-square images — which matches exactly the size-dependent, non-square-asymmetric residual error reported here, on top of the 0-1000 rescale from #16880.

    The PR author verified against a fine-tuned Qwen3-VL emitting per-word boxes at a fixed 1792×2560 input: llama.cpp before the fix matches transformers forced to align_corners=False bin-for-bin, and after the fix matches the real align_corners=True reference within bf16 noise — across ~220 boxes, fitting the pre-fix scale factor to 1.0126(x)/1.0134(y), matching the analytic prediction 1.0122/1.0149.

    Given this, would you consider adopting the align_corners fix from that PR into this fork's Qwen3VLChatHandler / vendored llama.cpp? Happy to test a build if useful.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions